GPT Image 2.5 misspells a brand name: spell it letter by letter
OpenAI's prompting docs say to quote the text and spell tricky words letter by letter. The prompt to write, and how to check the result on Sume's image API.

If GPT Image 2.5 misspells a brand name inside an image, put the exact wording in quotes, spell the tricky word out letter by letter in the prompt, and ask for no other text. Then read the result yourself, because OpenAI's own docs say the model can still struggle with text placement and clarity. On Sume the same prompt goes to openai/gpt-image-2.5 through POST /v1/images.
What OpenAI's docs say about text in images
The table below is limited to what OpenAI's current pages say. None of it is a guarantee of accuracy, and OpenAI's image guide still lists text placement and clarity as a limitation.
Source pages: Image prompting, the prompting guide in the OpenAI Cookbook and the Image generation guide.
| Topic | What the OpenAI page says |
|---|---|
| Exact wording | Put required wording in quotes and describe its position and typography. |
| Unusual words | Spell unusual words or brand names letter by letter when needed. |
| Extra text | Ask for no extra text, then check spelling and legibility in the output. |
| Small or dense text | Compare medium or high quality for small text, dense information or multiple fonts. |
| Known limit | The model can still struggle with precise text placement and clarity, even though it is significantly improved. |
A prompt that follows that advice
Name the surface, name the text once, give the spelling, and close the door on extra words. The invented word below is only there to show the pattern; swap in your own.
The request below keeps quality at high. On Sume, an omitted quality already defaults to high for this model, so setting it is a way to make the intent visible in your code.
curl -X POST "https://api.sume.com/v1/images" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"prompt": "Poster, flat navy background. Centered headline, exactly once: \"KOHLRABI\" (K-O-H-L-R-A-B-I), white bold sans-serif. No other text.",
"quality": "high",
"image_size": "1024x1536"
}'Check the output before it ships
Sume's image response is a hosted image URL plus a usage block. It has no field that reports the text found in the picture, so the spelling check is a human or a tool of your own. Open each result at the size it will be shown at, and compare the letters against your source string.
If a word is wrong, change one thing at a time: first the letter-by-letter spelling, then the quality tier, then the placement wording. OpenAI's iteration advice is to refine with small, single-change follow-ups and compare results before adding more instructions.
- Is every letter correct, including accents and capital letters?
- Does the text appear only where you asked, and only as many times as you asked? See why a headline appears twice.
- Is it legible at the real display size, not only zoomed in?
When to switch models
A failed spelling after two or three careful prompts is a sign to compare another model, not to keep retrying the same one. Failed Sume image generations are not billed, but a completed image with the wrong spelling is billed in full, so each retry on a finished image has a cost.
Sume's Image API docs say each model advertises the parameters it accepts, so read the catalog before you pin a model. For a side-by-side of catalog models that handle text, see the text-in-images comparison.
Sources
Related posts
More in Models
- Grok Imagine video in Sume: why it needs a start-frame image
Sume's Grok Imagine entry is image-to-video only: it blocks a submit without a start frame, tops out at 10 seconds, and sends no audio or aspect ratio.
- H3 Max 3D to Video: previs to photoreal on fal, not on Sume
fal's H3 Max 3D-to-Video turns a blockout render into photoreal video for $0.50 a request plus per second. Sume lists no such endpoint; what it offers instead.
- H3 Max Insert-Video: add a scene mid-clip on fal, and on Sume
fal's H3 Max Insert-Video adds 5 to 13 s to a source up to 60 s, billed on the new seconds only. Sume lists no such row; here is a trim, generate, join route.
- H3 Max Recast prompt: optional, and what Sume does with one
On Sume, H3 Max Recast runs without a prompt: the 1-4 photos say who to swap in. A prompt is allowed up to 2000 characters. Request body and limits below.
Written by Sume