Making Images and Diagrams Legible to Generative Systems
Most images on the web carry almost no retrievable information, not because images are inherently opaque to generative systems, but because almost nothing around them was built to make their content legible. Alt text gets filled in as an afterthought or skipped entirely. Filenames stay as the camera or export tool left them. Captions, when they exist at all, restate what a human viewer can already see rather than adding anything a system reading only the surrounding text would otherwise miss. This sub-chapter covers what actually changes that: treating an image's descriptive context as real content, not accessibility overhead.
- Alt text should describe what an image actually shows and means, not function as a keyword-stuffing opportunity
- Descriptive filenames are a small, frequently-skipped signal that costs nothing to get right
- Surrounding copy matters more for a diagram carrying unique information than for a purely illustrative photo
- Some diagrams express a relationship or process not stated anywhere else on the page, making the image the only source for that specific fact
- Captions serve a distinct purpose from alt text, primarily for human readers but read by systems too
- A practical audit approach can systematically surface which images in an existing library are genuinely legible and which are effectively invisible
Alt Text as Description, Not a Keyword Opportunity
Alt text exists to describe what an image actually shows, to a person who can’t see it and to a system reading the page’s underlying markup rather than rendering it visually. Treating alt text as a place to work in a target keyword phrase, regardless of whether that phrase accurately describes the image, produces text that fails both audiences at once: a screen reader user gets an inaccurate description, and a generative system reading that text gets a false signal about what the image actually contains.
The correct standard is simple to state and genuinely effortful to apply consistently across a large image library: write what the image shows, specifically enough that someone who can’t see it would understand its actual content. A product photo’s alt text should name the product and its relevant visible attributes. A diagram’s alt text should describe what relationship or process it depicts, not just label it generically as “diagram.” This isn’t a technical SEO checkbox distinct from the framework’s broader anti-manipulation stance; it’s the same honest-description principle applied to a different field on the page.
Descriptive Filenames: A Small, Frequently-Skipped Signal
A filename like IMG_4471.jpg carries no information. A filename like wide-toe-box-running-shoe-side-profile.jpg carries real, specific signal about what the file actually contains, independent of and in addition to whatever alt text ultimately gets applied to it.
This is a small fix with almost no cost, which is exactly why it’s worth naming directly rather than assuming it happens automatically. Export tools and content management systems default to generic, sequential filenames unless someone deliberately renames files before upload, and that deliberate step gets skipped constantly in practice, not because it’s difficult, but because it’s easy to forget when a dozen other things need attention during a content push. A quick, systematic pass renaming files descriptively before they go live costs a few minutes per image and adds a real, durable signal that persists for the life of the file.
Why Surrounding Copy Matters More for Diagrams Than Photos
A purely illustrative photo, one that shows a general scene already fully described in the surrounding text, carries relatively low retrieval stakes. If a system can’t process the image directly and the alt text is thin, little real information is lost, because the text around the photo already covers the same ground.
A diagram is a different case entirely. A diagram expressing a specific relationship, process, or comparison often exists precisely because that information is easier to communicate visually than in prose, which means the surrounding text frequently doesn’t restate it in full. This makes the quality of a diagram’s alt text and surrounding explanatory copy substantially more consequential than it is for a purely illustrative photo, since a thin description here risks losing information that exists nowhere else on the page.
When an Image Is the Only Source for a Specific Fact
The highest-stakes case is a diagram or chart that expresses a relationship, sequence, or comparison not stated anywhere else in the surrounding text. This happens more often than teams typically account for: a workflow diagram showing a process’s actual sequence, a comparison chart summarizing options a reader would otherwise have to piece together from scattered paragraphs, an architecture diagram showing how components connect.
In every one of these cases, the image isn’t illustrating already-explained content. It’s the primary source for that specific piece of information, and if a generative system can’t extract meaning from it, either through direct processing or through sufficiently detailed alt text and surrounding copy, that information is effectively unavailable to any answer the system generates. Identifying which images in a content library fall into this category, versus which are purely illustrative, is the single highest-value distinction to make when deciding where to invest description effort.
Captions as a Distinct Signal From Alt Text
A caption serves a genuinely different purpose from alt text, even though both describe an image in some sense. Alt text exists primarily for accessibility and for systems reading markup without rendering images visually. A caption exists primarily for a sighted human reader scanning a page, providing context, attribution, or a specific detail that adds to what the image visually shows rather than substituting for it.
Both get read by generative systems processing a page’s content, but they shouldn’t be treated as redundant or interchangeable. A caption that simply repeats the alt text wastes an opportunity to add genuinely distinct information, a source, a date, a specific detail about what’s shown that wouldn’t naturally belong in a pure accessibility description. Writing them as two distinct pieces of content, each doing its own job, produces more total retrievable signal than writing one and copying it into the other field.
A Practical Audit Approach for Existing Image Libraries
For a site with an existing library of images accumulated over time, a systematic audit is more productive than trying to fix everything at once. The practical approach starts by identifying which images fall into the high-stakes category covered above, diagrams and charts expressing information not stated elsewhere in the surrounding text, and prioritizing those first, since they represent the clearest cases of currently-inaccessible information.
From there, a broader pass through remaining images can check for the basics: generic or missing alt text, non-descriptive filenames, and captions that either don’t exist or simply duplicate the alt text without adding distinct value. This audit doesn’t need to happen all at once across an entire library; treating it as an ongoing practice applied to new content by default, with periodic passes through older content prioritized by which images carry the most unique information, is more sustainable than a one-time comprehensive fix that’s unlikely to actually get finished.
Making Images Carry Real, Retrievable Meaning
Michael Rubinstein has flagged image description as one of the most consistently under-invested areas across the sites he’s audited, not because it’s technically difficult, but because it sits in exactly the kind of low-visibility, easy-to-defer category that gets skipped when a content team is under time pressure, even though the actual cost of doing it well is genuinely small.
ScribePress generates descriptive alt text, filenames, and captions as a default part of its publishing pipeline, treating this as a required content field rather than an optional accessibility pass applied inconsistently after the fact.
Learn more about the work behind this framework at michael-rubinstein.com.
Frequently asked questions
Effective alt text describes what an image actually shows, specifically enough that someone who can't see it would understand its real content, rather than working in a target keyword regardless of accuracy. This serves both a screen reader user and a generative system reading the page's markup, since both depend on the alt text accurately reflecting what the image contains.
A descriptive filename is an additional, independent signal separate from alt text, and it costs almost nothing to implement correctly. Generic filenames like sequential camera exports carry no information on their own, while a descriptive filename adds a small but durable signal that persists for the life of the file, worth the minor effort of a deliberate renaming step before upload.
A purely illustrative photo typically shows something the surrounding text already fully describes, so a thin description loses little information. A diagram often exists specifically because its content is easier to communicate visually than in prose, meaning the surrounding text frequently doesn't restate what the diagram shows, which makes its alt text and surrounding copy substantially more consequential.
The highest-stakes images are diagrams and charts expressing a relationship, sequence, or comparison not stated anywhere else on the page, such as a workflow diagram or an architecture chart. These images function as the primary source for that specific information rather than illustrating already-explained content, making them the clearest cases where poor description results in genuinely lost information.
Alt text exists primarily for accessibility and for systems reading markup without rendering images visually, while a caption exists primarily for a sighted reader scanning the page and can add context an accessibility description wouldn't naturally include, like a source or a specific detail. Writing them as distinct content rather than duplicating one into the other produces more total retrievable signal.
Start with the highest-stakes category, diagrams and charts carrying information not stated elsewhere on the page, since these represent the clearest cases of currently-inaccessible content. From there, a broader ongoing pass through remaining images, prioritized by which carry the most unique information, is more sustainable than attempting a comprehensive one-time fix across an entire library.
Yes, the principle applies at any scale. A single diagram carrying unique information on an otherwise small site can be the only source for that specific fact just as much as one on a large site, and the effort required to describe a handful of images well is proportionally small regardless of how many images a site has in total.
To a meaningful degree, yes. If the surrounding text, alt text, and caption together describe what an image shows and means with real specificity, the information the image carries remains accessible to a system even if it can't process the visual content directly, which is precisely why this sub-chapter treats description quality as a genuine retrieval requirement rather than a secondary accessibility concern.
Put the framework to work
ScribePress
Turn GSO strategy into publish-ready content, straight into WordPress.
Visit ScribePress →