A 12-headline text rendering test for any image model

FLUX 3 and Ideogram 4.5 both claim better text. Run your own 12 headlines through the models Sume lists and review the outputs by eye, not by launch post.

4 min readSume
All posts

Test text rendering with your own copy: twelve headlines, the same prompt template, each listed model, and a person reading the outputs. Launch posts show best cases. Higgsfield describes FLUX 3 Image as having more accurate text and Ideogram 4.5 as doing text replacement that preserves layout (Higgsfield changelog, read 2026-10-02), but those are vendor-side descriptions, not your ads.

On Sume you can run the test across the text-capable models in the image catalog with one script and one key.

What goes in the headline set?

Choose headlines that stress the failure modes, not easy ones. The checks are visual, so keep a grid you can scan.

Headline set for the test, read 2026-10-02
TypeExampleWhat it stresses
ShortSALEBasic spelling
LongFree shipping on orders over 50 dollarsLine breaks
Numbers3 for 2, ends SundayDigits and punctuation
Brand nameAurelia Coffee RoastersUncommon spelling
Mixed caseNew: Cold-Brew ConcentrateHyphen and colon
AccentsCafe creme a la francaiseDiacritics if you add them

How do I run it?

Loop over models and headlines, send the same template, and print the result URLs. The model ids come from the OpenAPI schema; confirm them with the catalog call first. A 400 means that model rejected a field, so the script prints it and moves on.

import os
import requests

URL = "https://api.sume.com/v1/images"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MODELS = ["openai/gpt-image-2.5", "black-forest-labs/flux.2-flex",
          "ideogram/ideogram-v3"]
HEADLINES = ["SALE", "3 for 2, ends Sunday", "Aurelia Coffee Roasters"]

for model in MODELS:
    for text in HEADLINES:
        prompt = f"Cafe poster with the headline '{text}' in bold sans-serif"
        r = requests.post(URL, headers=HEADERS, timeout=90,
                          json={"model": model, "prompt": prompt})
        if r.status_code == 200:
            urls = [d["url"] for d in r.json().get("data", [])]
            print(model, repr(text), urls)
        else:
            print(model, repr(text), r.status_code, r.text[:120])

How do I score it?

Print the grid, mark each cell exact, one letter off, or wrong, and count per model. Repeat each cell once at a higher quality where the model lists it: Sume's docs say to escalate quality for finals and dense text. Image generation bills only on completion, so you pay for the images that come back, not for failures. The Image API docs list quality values per model.

Sources

Related posts

More in Models

All Models posts

Written by Sume