Translate text inside an image with GPT Image 2.5 edits
OpenAI's cookbook says to add 'do not change any other aspect of the image' when translating text in a picture. The edit call on Sume and what to verify.

To translate the text in an image with GPT Image 2.5, send the image as a reference, give the new wording in quotes, and add the line OpenAI's cookbook uses: do not change any other aspect of the image. On Sume that is POST /v1/images with openai/gpt-image-2.5, one reference, and aspect_ratio set to auto so the frame is not reshaped.
What OpenAI documents for this edit
The OpenAI Cookbook prompting guide lists translating text as an edit case and says to add that nothing else in the image should change. The image prompting page says to quote required wording, list what must stay the same (identity, geometry, layout, lighting, labels) and ask for no extra text.
| Step | What to write |
|---|---|
| Quote the new copy | Put the translated text in quotes. |
| Lock the rest | State that no other aspect of the image changes. |
| Name what must stay | List layout, lighting, labels and any logo. |
| Forbid extras | Ask for no extra text. |
| Verify | Check spelling and legibility in the output. |
The Sume call
A reference must be a public HTTPS URL. For edits, Sume's docs advise aspect_ratio: "auto" to match the reference, and note that omitting the field is not the same as auto.
curl -X POST "https://api.sume.com/v1/images" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-image-2.5",
"aspect_ratio": "auto",
"quality": "high",
"prompt": "Replace the sign text with \"OUVERT\". Do not change any other aspect of the image. No extra text.",
"input_references": [
{"type": "image_url",
"image_url": {"url": "https://example.com/storefront.png"}}
]
}'What can go wrong
The model redraws the whole image, so even a correct translation can shift a logo or a face. OpenAI's image guide says the model can struggle with text clarity and with keeping recurring brand elements consistent, which is the combination a translated sign or ad hits.
The OpenAI pages we read do not list which writing systems render reliably, so test each target language on a sample before a batch. Check accents, spacing and line breaks on every output.
Limit the area with a mask, if you can
Sume accepts an optional mask_url for GPT Image 2.5 edits. OpenAI says masking with GPT Image is prompt-based: the mask is guidance and the model may not follow its exact shape. Use it to narrow the edit, and still keep the unchanged-pixels instruction in the prompt. See how a mask edit works on Sume.
For a batch of translated variants, remember that each completed image is billed in full and failed ones are not, so run one language end to end before fanning out.
Sources
Related posts
More in Use cases
- Halloween product teaser: a 15-second vertical clip from one still
Make a 15-second Halloween teaser for a product with Seedance 2.5 on Sume: first-frame still, 9:16 at 720p, native audio, and how to read the price first.
- Halloween narrator voiceover with spooky music for a video
Make a Halloween story video audio track: a slow narrator from Sume TTS, an instrumental horror bed from the music router, mixed with ducking in Timeline 1.0.
- Higgsfield's Seedance outage on Sept 30: slow vs failed jobs on Sume
Higgsfield said Seedance 2.0 failed more and 2.5 ran slow on Sept 30, now fixed. On Sume, a slow job is queued or processing; rerun only after failed.
- Check 200 holiday clips before upload with a free probe-only inspect
A video-inspect call with frames false is a probe with no stills and no charge. Use it as a pre-upload audio gate, then pull stills only for clips you doubt.
Written by Sume