Chapter 12 · Pillar

Chapter 12: Video, Visual, and Multimodal GSO

Eleven chapters of this framework have been built around text-first assumptions, because that's where GSO needed to start. Generative systems don't stop at text, though, and neither does this chapter. Video, images, audio, and the platforms multimodal content actually lives on are all genuine source material a system can retrieve from and reason about, provided they're built to be retrievable rather than treated as decoration around the real content. This chapter establishes that principle once, points to where video's full depth already lives, and gives images, audio, and platform ecosystems the dedicated treatment they don't yet have anywhere else on this site.

Key takeaways
  • Multimedia content is genuine source material a generative system retrieves from and reasons about, not supplementary branding
  • Video already has a complete, dedicated treatment at the Video GSO flagship piece; this chapter points there rather than repeating it at lower depth
  • Multimedia pages apply the same spoke discipline established in Chapter 8.3, not a new, seventh functional page type
  • Images, audio, and platform ecosystems each get dedicated coverage here, since no existing resource on this site covers them yet
  • Multimodal source coherence, the requirement that the same entity tell the same story across every medium, is this chapter's throughline and closing synthesis
  • This chapter resolves a real gap between the Video GSO piece's own coherence argument and Chapter 6.3's source coherence work, connecting the two explicitly

Why This Chapter Exists, and Why It’s Shorter Than It Could Be

Generative systems process more than written pages, and a brand that treats its video, image, and audio content as branding rather than source material is leaving real, retrievable information unreachable. That’s the case this chapter makes, and it’s a genuine gap in everything the prior eleven chapters have covered, which assumed text as the default medium throughout.

What this chapter deliberately doesn’t do is repeat depth that already exists elsewhere on this site. Video has a complete, thorough treatment at the Video GSO flagship piece, and rewriting that ground at doctrine length would produce a worse, shorter version of something that already works. This chapter is built around that decision rather than around it.

Multimodal Content as Source Material

The chapter opens by establishing its core principle: multimedia content is genuine source material, not decoration, and video demonstrates this most clearly of any medium. Rather than re-deriving video’s specific mechanics, this sub-chapter states the principle, defers depth to the Video GSO piece explicitly, and makes the architectural decision that multimedia pages apply spoke discipline rather than requiring a new functional page type.

Chapter 12.1 covers this in full, including how the retrieval pipeline from Chapter 3 extends to non-text input.

Images and Diagrams

Alt text, filenames, and captions are usually treated as accessibility overhead rather than real content, and that treatment is exactly what makes most images retrievably invisible. This sub-chapter distinguishes purely illustrative images from diagrams carrying information stated nowhere else on a page, and covers what makes each genuinely legible.

Chapter 12.2 covers this ground, which no existing resource on this site addresses.

Audio and Podcasts

Audio shares real ground with video, transcripts matter for both, but lacks video’s visual chapter markers, which makes show notes, timestamps, and clear speaker attribution carry more structural weight than they do for other media.

Chapter 12.3 covers what’s specifically different about audio once the ground it shares with video is acknowledged rather than repeated.

Platform Ecosystems

A significant share of a brand’s multimodal footprint lives on third-party platforms, YouTube, LinkedIn, podcast hosts, review sites, rather than its own domain. This is a different axis from the prior three sub-chapters: not what type of content this is, but where it lives, and what has to stay consistent regardless of platform.

Chapter 12.4 covers this, including why platform presence itself functions as an external validation signal, per Chapter 10.4.

Multimodal Coherence

Every prior sub-chapter in this chapter serves one requirement: the same entity needs to tell the same story across every medium it publishes in. This isn’t a new concept, it’s Chapter 6.3’s source coherence work, extended with media type as another axis of “surface.”

Chapter 12.5 makes this connection explicit and resolves a real gap between this framework’s own doctrine and the Video GSO piece’s independent version of the same argument, closing the chapter by tying every sub-chapter’s guidance to this single requirement.

From Doctrine to Operation

Everything covered across this chapter still needs to actually get built, audited, and maintained as part of a real content operation, not just understood as principle. Chapter 13 picks up directly from here, covering how GSO, including the multimodal work this chapter establishes, operates as a repeatable system rather than a one-time project.

Establishing the Principle, Not Repeating What Already Works

Michael Rubinstein built this chapter around a discipline the framework asks of every practitioner it teaches: when genuine depth already exists somewhere, point to it rather than compete with a shorter version of the same thing. Video already had that depth. This chapter’s job was recognizing that, stating the principle it demonstrates, and doing equal-quality work on the media types that didn’t have anywhere else to live.

ScribePress applies this same multimodal source-material discipline across every format it produces, treating video, image, and audio content as retrievable source material from the point of production rather than branding retrofitted with better metadata after the fact.

Learn more about the work behind this framework at michael-rubinstein.com.

Frequently asked questions

The Video GSO flagship piece already covers video's full mechanics, transcripts, schema, chapter markers, in complete depth, and repeating that ground at doctrine length would only produce a shorter, worse version of something that already exists and works well. This chapter states the underlying principle video demonstrates and points to that piece directly rather than competing with it.

Multimedia content, video, images, audio, is genuine source material a generative system retrieves from and reasons about, not supplementary branding. Content built so the media itself is retrievable, transcripts, detailed alt text, structured show notes, produces a fundamentally different outcome than content that merely happens to include media alongside text that does the real explanatory work.

No. This chapter makes an explicit architectural decision that multimedia pages apply the same spoke discipline covered in Chapter 8.3 to a different medium, rather than requiring a seventh functional page type. This keeps the site's overall architecture unified rather than fragmenting it into medium-specific categories.

It's the requirement that the same entity, expertise, claims, and positioning stay consistent across every medium a brand publishes in, directly extending the source coherence work established in Chapter 6.3 with media type as another axis. This chapter's closing sub-chapter, 12.5, covers this in full and ties every other sub-chapter's guidance to this single requirement.

Platform ecosystems represent a different axis entirely, where content lives rather than what type it is. A significant share of a brand's multimodal footprint exists on third-party platforms rather than its own domain, and platform presence itself functions as an external validation signal, connecting directly to the trust-building work in Chapter 10.4.

No, deliberately. Multimodal AI capability changes quickly enough that a specific claim about a named model's current ability risks becoming outdated within months. Every principle in this chapter holds regardless of current model capability, since well-structured multimedia content is correct practice whether a system processes the media directly or relies on its transcript, description, or captions.

Chapter 13 picks up directly from here, covering how the principles established across this framework, including this chapter's multimodal work, operate as part of a repeatable operating system rather than one-time guidance applied once and left unmaintained.

Put the framework to work

ScribePress

Turn GSO strategy into publish-ready content, straight into WordPress.

Visit ScribePress →
WhatsApp