Chapter 12.1 · Spoke

Treating Multimodal Content as Source Material, Not Branding

A video embedded on a page is either two things: a piece of source material a generative system can retrieve from and reason about, or a decorative asset that happens to sit near the text doing the actual work. Most sites, without deciding to, end up with the second. This sub-chapter makes the case for the first, and it does so without repeating ground this framework has already covered in exhaustive depth elsewhere. Video specifically already has a full, dedicated treatment at [gsoguide.online/video-gso/](/video-gso/), and this page points there deliberately rather than compressing that depth into something shorter and worse. What follows is the principle underneath all of it, stated once, and applied to the media types that genuinely need their own room: images, audio, and the platforms multimodal content actually lives on.

Key takeaways
  • Video, audio, and image content are genuine source material a generative system retrieves from and reasons about, not supplementary branding assets
  • Full video-specific mechanics, transcripts, schema, chapter markers, already exist in complete depth at the Video GSO flagship piece and are not repeated here
  • The retrieval and synthesis pipeline covered in Chapter 3 extends to non-text input; it doesn't require a separate mechanism for each modality
  • Multimedia pages apply the same spoke discipline established in Chapter 8.3, not a new, seventh functional page type
  • Multimodal source coherence, the throughline for this entire chapter, gets its full treatment in Chapter 12.5
  • Images, audio, and platform ecosystems each still need dedicated coverage of their own, which the rest of this chapter provides

How the Retrieval Pipeline Extends to Non-Text Input

The retrieval, evaluation, fragment selection, and synthesis pipeline covered in Chapter 3 was described in text-first terms because that’s where the framework needed to start. Nothing about the mechanism itself is text-exclusive. A system retrieving material to answer a prompt can draw on a video transcript, an image’s alt text and caption, or a podcast’s show notes the same way it draws on a written paragraph, provided that material is structured to be reached and understood in the first place.

This isn’t a new pipeline for multimodal content. It’s the same one, with more input types available to it. The practical implication is direct: everything this framework has already established about crawlability, rendering, and fragment-level clarity applies to a transcript or a caption exactly as it applies to a paragraph of body text, which is why this chapter leans so heavily on citing earlier chapters rather than re-deriving their mechanics for each new medium.

Branding vs. Source Material

Most video, image, and audio content on the web was built to support a page, not to be a page’s actual substance. A hero video plays behind a headline. A product photo sits next to a description that does the real explaining. A podcast episode gets embedded with a two-sentence summary and nothing else. In every one of these cases, the multimedia asset is decoration, and the text around it is what actually carries retrievable meaning.

The alternative is building multimedia content so the media itself is the retrievable substance: a video whose transcript states the same claims a written page would, an image whose surrounding context explains what it shows clearly enough that the image’s information doesn’t disappear if a system can’t process the pixels directly, a podcast whose show notes function as a real written record rather than a marketing blurb. This is the distinction this entire chapter is built around, and it’s a decision made at the content-planning stage, not something fixed after the fact with better alt text alone.

Video as the Proof Case, and Where Its Full Depth Actually Lives

Video is the clearest demonstration of the branding-versus-source-material distinction, and it’s also the medium this framework has already covered most thoroughly. The full mechanics, transcript structure as the retrieval foundation, VideoObject schema properties down to chapter-level markup, chapter titles functioning as declarations of what questions a video actually answers, are documented in complete depth at gsoguide.online/video-gso/, and repeating that material here at doctrine length would only produce a worse, shorter version of something that already exists and works.

This is a deliberate architectural decision, not an omission. A doctrine chapter’s job is establishing principles; a flagship research piece’s job is going deep on one of them. Video already has its deep treatment. What this chapter needs to do instead is establish the principle clearly enough that a reader understands why video was treated this way, and then do the equivalent work for the media types that don’t yet have that depth anywhere on this site.

The Architectural Decision: Multimedia Pages Are Spokes

Chapter 8.4 named six functional page types: pillar, spoke, glossary, comparison, FAQ, and evidence page. No video- or audio-specific type appears among them, and that absence is intentional rather than an oversight this chapter needs to correct.

A video page, an image-heavy explainer, or a podcast episode page is applying the same spoke discipline established in Chapter 8.3 to a different medium: it should resolve one specific intent cluster completely, remain self-contained without requiring outside context to make sense, and be built from clearly structured, extractable pieces, whether those pieces are paragraphs or transcript segments. Introducing a seventh functional type for multimedia content would fragment the site’s architecture without adding anything a properly built spoke doesn’t already provide. This is a deliberate choice worth stating plainly rather than leaving a reader to wonder why video didn’t get its own category in Chapter 8.

What Multimodal Source Coherence Means, at a Preview Level

Every principle in this chapter ultimately serves one requirement: the same entity, expertise, claims, and positioning need to be reinforced consistently across every medium a brand publishes in, not just across written pages. A claim stated one way in text and a subtly different way in a still-live video isn’t two separate pieces of content aging independently. It’s one inconsistency a generative system can encounter from either direction.

This is Chapter 12.5’s full subject, and it deserves the complete treatment that closing sub-chapter gives it rather than a partial version here. What matters at this point is understanding that everything covered across the rest of this chapter, images, audio, platform ecosystems, is ultimately in service of this one coherence requirement, not a scattered set of unrelated medium-specific tips.

Why Images, Audio, and Platforms Still Need Their Own Chapters

Video got the deferral treatment because a thorough resource already exists for it. Images, audio, and the platform ecosystems multimodal content actually lives on don’t have that resource anywhere on this site yet, which is exactly why they still get dedicated sub-chapters rather than the same brief treatment.

Chapter 12.2 covers what makes an image or diagram genuinely interpretable rather than decorative. Chapter 12.3 covers what’s specifically different about audio-only content once the transcript principles it shares with video are accounted for. Chapter 12.4 covers an entirely different axis: not what type of content this is, but where it lives, since a great deal of multimodal content sits on third-party platforms rather than a brand’s own domain. Each of these covers real, currently underserved ground, which is the same standard video was held to before this framework decided its depth already existed elsewhere.

Establishing the Principle, Then Pointing to Where Depth Already Lives

Michael Rubinstein has treated this chapter’s scope as a direct test of whether this framework practices its own internal-linking discipline from Chapter 8.6 or just states it: a resource that already exists at genuine depth doesn’t need a shorter, worse doctrine-chapter version competing with it, it needs an honest pointer and the doctrine-level principle that explains why the pointer is there.

ScribePress treats multimedia content with the same source-material discipline this page describes by default, structuring video, image, and audio assets to be genuinely retrievable rather than decorative from the point of production, not retrofitted afterward.

Learn more about the work behind this framework at michael-rubinstein.com.

Frequently asked questions

Video already has a complete, dedicated treatment at the Video GSO flagship piece, covering transcript structure, schema implementation, and chapter markers in depth this doctrine chapter would only repeat at lower quality. Images, audio, and platform ecosystems don't have an equivalent resource anywhere on this site yet, so they receive full dedicated sub-chapters here instead.

The same pipeline, retrieval, evaluation, fragment selection, and synthesis, applies regardless of input type. A system can draw on a video transcript, an image's caption, or a podcast's show notes the same way it draws on written text, provided that material is structured to be reached and understood, which is why the crawlability and clarity principles established for text apply equally to multimedia content.

Branding treatment means the media supports a page while the surrounding text does the real explanatory work, a hero video playing behind a headline, for instance. Source-material treatment means the media itself carries retrievable substance, a video transcript stating the same claims a written page would, an image with context specific enough that its information doesn't disappear if a system can't process the pixels directly.

No. This is a deliberate architectural decision: multimedia pages apply the same spoke discipline established in Chapter 8.3 to a different medium, rather than requiring a seventh functional type. A video or podcast episode page should resolve one specific intent cluster completely and remain self-contained, the same standard any other spoke is held to.

It's the requirement that the same entity, expertise, claims, and positioning stay consistent across every medium a brand publishes in, not just across written pages. This chapter's full treatment of the concept lives in Chapter 12.5; every other sub-chapter in this chapter ultimately serves this single coherence requirement.

The Video GSO flagship piece at gsoguide.online/video-gso/ covers this in complete depth, including transcript structure, VideoObject schema properties, and chapter markers as query-coverage signals. This chapter deliberately doesn't repeat that material and points there directly instead.

Multimodal AI capability is changing fast enough that a specific claim about a named model's current ability risks being wrong or outdated within months. Every principle in this chapter is framed to hold regardless of current model capability, since a well-structured video, image, or audio asset is correct practice whether a system processes the media directly, reads its transcript, or both.

No. The branding-versus-source-material distinction applies at any scale, a single podcast episode or a single explanatory diagram can be built either way. The principle doesn't require a large content library to matter; it requires a deliberate choice at the point a piece of multimedia content gets planned, regardless of how much content a team ultimately produces.

Put the framework to work

ScribePress

Turn GSO strategy into publish-ready content, straight into WordPress.

Visit ScribePress →
WhatsApp