AI image text: what each vendor claims and how to test it
OpenAI admits text placement can fail, Google promises legible stylized text. Four vendor pages, what they actually claim, and a test harness to run on Sume.

Only two of the four vendor pages make a statement about text in images, and they point in different directions. OpenAI says GPT Image 2.5 is better but "can still struggle with precise text placement and clarity." Google says its Nano Banana models can generate "legible, stylized text for infographics, menus, diagrams, and marketing assets." The xAI and Black Forest Labs pages read for this post make no text claim.
That is a reason to test, not to trust a headline. The harness below sends the same quoted copy to every text-capable row in the Sume catalog and gives you URLs to compare by eye.
What exactly did each page say?
Quotes are from the pages as read on 2026-10-02. A missing claim is not a failure; it means the page was silent.
A vendor sentence is also a snapshot. Models change, and a page that is silent today may add a claim next month. Put the date next to every number you keep, as the table caption does, and rerun the harness when a new model id appears in your catalog.
| Vendor page | What it says about text | Take-away |
|---|---|---|
| OpenAI image guide | Improved, but can struggle with placement and clarity | Check placement and spelling every time |
| Google Gemini image guide | Legible, stylized text for infographics, menus, diagrams | A stronger claim; still verify |
| xAI Imagine guide | No text-rendering statement found | Unknown, test it |
| Black Forest Labs docs home | No text-rendering detail on that page | Unknown, test it |
How do you build a fair test?
Use identical copy and layout wording, a few lengths, and more than one sample, because text errors are random.
- Three strings: one word, a short headline, a two-line price tag.
- Quote each string and say how many times it appears.
- Fix the aspect ratio so the layout is comparable.
- Generate at least three images per model before judging.
- Count wrong letters, doubled words and misplaced lines, not overall beauty.
The harness
It lists catalog rows that accept a prompt, filters by name fragments you choose, and prints one URL per model and string. Edit WANT to match the models you care about; the ids come from your own catalog, not from this post.
import os, requests
B = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
WANT = ("gpt-image-2.5", "banana", "grok-imagine-image", "flux", "seedream")
COPY = ['Poster with the text "OPEN LATE" once, centred, bold sans-serif',
'Menu card titled "Lunch 11-3" with two lines below it']
models = [m["id"] for m in requests.get(B + "/v1/images/models", headers=H).json()["data"]
if any(w in m["id"] for w in WANT)]
for mid in models:
for prompt in COPY:
r = requests.post(B + "/v1/images", headers=H, timeout=90,
json={"model": mid, "prompt": prompt})
out = r.json()["data"][0]["url"] if r.status_code == 200 else f"HTTP {r.status_code}"
print(mid, "|", prompt[:30], "|", out)What does a pass look like?
A pass is every character correct, the copy appearing exactly once, and the text sitting where you asked. If a model fails twice on a short string, stop. Compose the text outside the model and use the model for the picture, as the earlier text posts suggest. Rows that returned 202 need a job read; the harness prints the status code for them.
Keep the output. Save each URL, the prompt and the model id in a sheet, then score the rows by hand. Ten minutes of scoring beats an afternoon of arguing from benchmarks, and the sheet becomes the record you can show a client when they ask why you chose one model for their packaging.
Sources
Related posts
More in Comparisons
- AI video cost per minute: Veo 3.1 vs the Sume catalog
One minute of generated footage costs $4.80 to $24 on Veo 3.1 at list prices and $0.75 to $15 on the Sume video catalog. Per-second rates, dated 2026-10-01.
- AI/ML API vs Sume: OpenAI-style base URL or media jobs
AI/ML API gives an OpenAI-compatible base URL across text, image, video and music. Sume is media jobs, and its agent endpoint is async, not choices[].
- Alibaba Wan 3.0 on OpenRouter vs Sume: id, seed and aspect ratios
OpenRouter lists alibaba/wan-3.0 with seed and 2 to 30 seconds. Sume lists wan-3.0 with the same range, no seed, and its own billing. What to change.
- Amazon Ads Agent creative tools vs Sume for off-Amazon video
Amazon's unBoxed 2026 Ads Agent now makes TV-quality video from product pages. What it covers, and how Sume fills the TikTok, Meta and Shopify side.
Written by Sume