A 12-headline text rendering test for any image model
FLUX 3 and Ideogram 4.5 both claim better text. Run your own 12 headlines through the models Sume lists and review the outputs by eye, not by launch post.

Test text rendering with your own copy: twelve headlines, the same prompt template, each listed model, and a person reading the outputs. Launch posts show best cases. Higgsfield describes FLUX 3 Image as having more accurate text and Ideogram 4.5 as doing text replacement that preserves layout (Higgsfield changelog, read 2026-10-02), but those are vendor-side descriptions, not your ads.
On Sume you can run the test across the text-capable models in the image catalog with one script and one key.
What goes in the headline set?
Choose headlines that stress the failure modes, not easy ones. The checks are visual, so keep a grid you can scan.
| Type | Example | What it stresses |
|---|---|---|
| Short | SALE | Basic spelling |
| Long | Free shipping on orders over 50 dollars | Line breaks |
| Numbers | 3 for 2, ends Sunday | Digits and punctuation |
| Brand name | Aurelia Coffee Roasters | Uncommon spelling |
| Mixed case | New: Cold-Brew Concentrate | Hyphen and colon |
| Accents | Cafe creme a la francaise | Diacritics if you add them |
How do I run it?
Loop over models and headlines, send the same template, and print the result URLs. The model ids come from the OpenAPI schema; confirm them with the catalog call first. A 400 means that model rejected a field, so the script prints it and moves on.
import os
import requests
URL = "https://api.sume.com/v1/images"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MODELS = ["openai/gpt-image-2.5", "black-forest-labs/flux.2-flex",
"ideogram/ideogram-v3"]
HEADLINES = ["SALE", "3 for 2, ends Sunday", "Aurelia Coffee Roasters"]
for model in MODELS:
for text in HEADLINES:
prompt = f"Cafe poster with the headline '{text}' in bold sans-serif"
r = requests.post(URL, headers=HEADERS, timeout=90,
json={"model": model, "prompt": prompt})
if r.status_code == 200:
urls = [d["url"] for d in r.json().get("data", [])]
print(model, repr(text), urls)
else:
print(model, repr(text), r.status_code, r.text[:120])How do I score it?
Print the grid, mark each cell exact, one letter off, or wrong, and count per model. Repeat each cell once at a higher quality where the model lists it: Sume's docs say to escalate quality for finals and dense text. Image generation bills only on completion, so you pay for the images that come back, not for failures. The Image API docs list quality values per model.
Sources
Related posts
More in Models
- Veo 3.1 and Veo 3.1 Fast previews end October 22: what replaces them
Google's deprecations page sets October 22, 2026 as the shutdown date for the Veo 3.1 and 3.1 Fast previews. Dates, the replacement id and what Sume lists.
- Veo seed doesn't make output repeatable; Sume rejects seed
Google says Veo's seed only slightly improves determinism. Sume's v1 video models report seed false and reject the field; keep the output file to get a repeat.
- Video model leaderboard rank: how to turn Elo into an API choice
Hedra's board puts Seedance 2.0 at Elo 1,225 (#2) and HappyHorse 1.1 at 1,149 (#4). What a rank tells an API buyer and what to test instead.
- Which AI video model gives 1080p on Sume, and which stop at 768p?
Sume lists 1080p for Seedance, Wan 3.0, Kling 3.0 and Auto; MiniMax H3 is native 480p or 768p in the panel. Resolution table by model, with the API check.
Written by Sume