Can AI image models spell a headline correctly? A pre-ship text check
Vendors say image models can render text but still miss. Use a short script to compare the wording you asked for with what a person typed back from the image.

Sometimes, and you cannot tell by looking at the model name. OpenAI's image guide (read 2026-10-06) says its image models can still struggle with precise text placement and clarity. Google's Gemini image docs (read 2026-10-06) describe legible text for things like infographics, menus and diagrams. Both are claims about capability, not guarantees about your headline.
So treat text in an image as something you verify, every time, before the image leaves your team.
A check that takes a minute
Have a person read the finished image and type exactly what they see into a text file. The script compares that against the wording you asked for and fails when they differ. It is deliberately dull: no scoring model, no threshold you cannot explain.
import difflib, sys
def check(expected: str, seen: str) -> int:
a, b = expected.strip(), seen.strip()
if a == b:
print("exact match")
return 0
ratio = difflib.SequenceMatcher(None, a, b).ratio()
print(f"MISMATCH ratio={ratio:.2f}")
for op, i1, i2, j1, j2 in difflib.SequenceMatcher(None, a, b).get_opcodes():
if op != "equal":
print(f" {op}: asked {a[i1:i2]!r} saw {b[j1:j2]!r}")
return 1
if __name__ == "__main__":
if len(sys.argv) != 3:
raise SystemExit("usage: check.py EXPECTED SEEN")
sys.exit(check(sys.argv[1], sys.argv[2]))Where Sume helps and where it does not
| Question | Answer | Source |
|---|---|---|
| Does Sume guarantee spelling? | No. No image row advertises it | Image API docs |
| Which Sume rows take no references (text to image only)? | Imagen 4 Fast and Ultra, Recraft V4, Qwen Image Max, Soul | Image API docs |
| Can I fix one word in an existing image? | Yes on rows that take references: send the image and ask for the one change | Image API docs |
| Does the image need alt text? | If the text is not elsewhere on the page, put it in the alt | W3C WAI |
A workable loop
- Generate the layout once at low or medium quality and approve the design, not the words.
- Run an edit pass that changes only the headline, and keep the other elements in the instruction as unchanged.
- Run the check. If it fails, edit again from the last good file, not from scratch.
- Cap the attempts at three. If the word still fails, set the headline in your design tool instead.
Do not forget the alt text
W3C's images tutorial (read 2026-10-06) says that when an image contains text, the text needs to be in the alt attribute unless it appears elsewhere on the page. Store the approved wording with the file so the page author does not have to retype it from the picture.
Sources
Related posts
More in Use cases
- Coffee shop loyalty card art with an API: stamps drawn in code
Generate 3:2 loyalty card art on Sume with empty stamp spaces, then draw each customer's stamps in code so the count is always right and the art is reused.
- Convert a YouTube Shorts playlist into a show with Add show features
In YouTube Studio, Options on a playlist offers Add show features, and every video lands in Season 1. What to fix first and how to prepare Shorts for it.
- Criteo Retail Media onsite video: 720p, 15 s, a Sume avatar clip
Criteo onsite video asks for MP4 at 720p, 15 s recommended and 30 s max, in 16:9, 1:1 or 9:16. Sume avatar output is 720p, so script to 15 seconds.
- Narration budget for a daily short-video channel: 30 videos a month
Thirty 750-character narrations a month cost $1.20 on Sume TTS with per-job rounding. The same characters are about $0.50 on MAI-Voice-2.1 and $0.34 on Flash.
Written by Sume