AI image text: what each vendor claims and how to test it

OpenAI admits text placement can fail, Google promises legible stylized text. Four vendor pages, what they actually claim, and a test harness to run on Sume.

5 min readSume
All posts

Only two of the four vendor pages make a statement about text in images, and they point in different directions. OpenAI says GPT Image 2.5 is better but "can still struggle with precise text placement and clarity." Google says its Nano Banana models can generate "legible, stylized text for infographics, menus, diagrams, and marketing assets." The xAI and Black Forest Labs pages read for this post make no text claim.

That is a reason to test, not to trust a headline. The harness below sends the same quoted copy to every text-capable row in the Sume catalog and gives you URLs to compare by eye.

What exactly did each page say?

Quotes are from the pages as read on 2026-10-02. A missing claim is not a failure; it means the page was silent.

A vendor sentence is also a snapshot. Models change, and a page that is silent today may add a claim next month. Put the date next to every number you keep, as the table caption does, and rerun the harness when a new model id appears in your catalog.

Text-in-image statements by vendor page (read 2026-10-02)
Vendor pageWhat it says about textTake-away
OpenAI image guideImproved, but can struggle with placement and clarityCheck placement and spelling every time
Google Gemini image guideLegible, stylized text for infographics, menus, diagramsA stronger claim; still verify
xAI Imagine guideNo text-rendering statement foundUnknown, test it
Black Forest Labs docs homeNo text-rendering detail on that pageUnknown, test it

How do you build a fair test?

Use identical copy and layout wording, a few lengths, and more than one sample, because text errors are random.

  • Three strings: one word, a short headline, a two-line price tag.
  • Quote each string and say how many times it appears.
  • Fix the aspect ratio so the layout is comparable.
  • Generate at least three images per model before judging.
  • Count wrong letters, doubled words and misplaced lines, not overall beauty.

The harness

It lists catalog rows that accept a prompt, filters by name fragments you choose, and prints one URL per model and string. Edit WANT to match the models you care about; the ids come from your own catalog, not from this post.

import os, requests
B = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
WANT = ("gpt-image-2.5", "banana", "grok-imagine-image", "flux", "seedream")
COPY = ['Poster with the text "OPEN LATE" once, centred, bold sans-serif',
        'Menu card titled "Lunch 11-3" with two lines below it']
models = [m["id"] for m in requests.get(B + "/v1/images/models", headers=H).json()["data"]
          if any(w in m["id"] for w in WANT)]
for mid in models:
    for prompt in COPY:
        r = requests.post(B + "/v1/images", headers=H, timeout=90,
                          json={"model": mid, "prompt": prompt})
        out = r.json()["data"][0]["url"] if r.status_code == 200 else f"HTTP {r.status_code}"
        print(mid, "|", prompt[:30], "|", out)

What does a pass look like?

A pass is every character correct, the copy appearing exactly once, and the text sitting where you asked. If a model fails twice on a short string, stop. Compose the text outside the model and use the model for the picture, as the earlier text posts suggest. Rows that returned 202 need a job read; the harness prints the status code for them.

Keep the output. Save each URL, the prompt and the model id in a sheet, then score the rows by hand. Ten minutes of scoring beats an afternoon of arguing from benchmarks, and the sheet becomes the record you can show a client when they ask why you chose one model for their packaging.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume