GPT Image 2.5 text in the wrong place: fix it with a mask edit
OpenAI says GPT Image 2.5 can still misplace text. Mask only the text zone and re-render that region through Sume's mask_url edit instead of the whole image.

When GPT Image 2.5 puts a headline in the wrong spot, do not regenerate the whole picture. Make a mask that covers only the text area and run one edit with mask_url; the rest of the image has the best chance of staying as it was.
OpenAI's image generation guide is direct about the limit: although text is "significantly improved, the model can still struggle with precise text placement and clarity." That matches what the Sume Image API offers for fixing it: an optional mask_url on ChatGPT Image 2.5 edits.
What does Sume accept for this edit?
The Image API docs say openai/gpt-image-2.5 (Flare) and openai/gpt-image-2.5-sunburst both support up to 16 image references, an optional mask_url, and background: auto|transparent|opaque. The mask URL must be public HTTPS, like every reference.
OpenAI's guide says masks identify the areas to replace, and that the image and mask must be the same format and size, under 50MB, with an alpha channel. Build your mask to that rule before uploading it somewhere public.
How do you write the prompt?
Name the text exactly, say where it goes, and say what must not change. Keep the quoted copy short, because longer strings are where placement and spelling slip first.
- Quote the string:
Headline reads "SUMMER SALE"in straight quotes. - Describe the zone:
top third, centred, white sans-serif on the dark sky. - Add the preserve line:
Keep the product, background and lighting unchanged. - State the count: say the text appears once.
A request you can send
Replace the two URLs with your own hosted image and mask. The model id comes straight from the docs.
Two practical notes on the call. First, keep aspect_ratio at auto so the edit keeps the original's shape; the Image API docs say omitting the field is not the same as auto. Second, pick the quality deliberately. Omitted quality defaults to high on Sume, and auto reserves the max amount, so a text-only touch-up does not need either. Try medium first and move up only if the letters are soft.
import os, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
body = {
"model": "openai/gpt-image-2.5",
"prompt": 'Headline reads "SUMMER SALE" once, centred in the top third, '
"white sans-serif. Keep the product, background and light unchanged.",
"input_references": [{"type": "image_url",
"image_url": {"url": "https://example.com/ad.png"}}],
"mask_url": "https://example.com/ad-text-zone-mask.png",
"aspect_ratio": "auto",
"quality": "high",
}
r = requests.post("https://api.sume.com/v1/images", headers=H, json=body, timeout=60)
print(r.status_code, r.json())What should you expect from the result?
A mask is guidance, not a hard clip. Our earlier post on edits that leaked outside the mask explains why, and the pixel-diff post shows how to measure it. Plan on checking the result once rather than trusting it.
If the text is still wrong after two attempts, stop paying for retries. Drop the text from the image, render it in your design tool, and use the model for the picture only. That is the honest answer when exact typography matters, for example a legal line or a price.
A final tip on building the mask. Make the white (edit) zone slightly larger than the old text, not tight around it. A tight mask leaves ghost remnants of the previous letters at the edges, while a generous one gives the model room to repaint the background behind the new words. Remember the OpenAI rules from the checklist post: same size and format as the image, with an alpha channel.
| Approach | What can change | Best when |
|---|---|---|
| Regenerate the whole image | Everything | The layout itself is wrong |
Masked edit with mask_url | Mostly the masked zone | Only the text is wrong |
| Typeset outside the model | Nothing in the picture | Exact copy is mandatory |
Sources
Related posts
More in Models
- grok-4.5 as a Sume model id runs GPT-6 Sol, not Grok: the alias list
Sume treats grok-4.5, composer-2.5, composer-2.5-fast and sume-1.0-fast as old Cursor-era ids that run GPT-6 Sol. How Grok 4.7 is gated and how grok-4.6 moves.
- Grok Imagine video in Sume: why it needs a start-frame image
Sume's Grok Imagine entry is image-to-video only: it blocks a submit without a start frame, tops out at 10 seconds, and sends no audio or aspect ratio.
- H3 Max 3D to Video: previs to photoreal on fal, not on Sume
fal's H3 Max 3D-to-Video turns a blockout render into photoreal video for $0.50 a request plus per second. Sume lists no such endpoint; what it offers instead.
- H3 Max Insert-Video: add a scene mid-clip on fal, and on Sume
fal's H3 Max Insert-Video adds 5 to 13 s to a source up to 60 s, billed on the new seconds only. Sume lists no such row; here is a trim, generate, join route.
Written by Sume