AI change text in image: an edit call with a reference
To change wording inside a picture, send it as an input_references image with a prompt giving the new text. Sume lists edit-capable models.

Send the picture as an input_references image to POST /v1/images with a prompt that quotes the old wording and the new wording, and use aspect_ratio: "auto" to keep its shape. The example uses google/nano-banana-2, which is edit-capable and lists auto; Sume also lists Ideogram V3 as edit-capable, but its ratio list has no auto. Results can change more than the text, so check them; for exact wording, overlay the text in your own tool.
Request rules are from Image generation and editing, read 2026-09-30. Ideogram's 4.5 page lists text modification among its edit uses; Sume's catalog lists V3, not 4.5.
What does the call look like?
Name the text to replace and the text to use, and tell the model to leave the rest alone.
const res = await fetch("https://api.sume.com/v1/images", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "google/nano-banana-2",
prompt: 'Change the sign text from "OPEN" to "SALE". Keep everything else unchanged.',
aspect_ratio: "auto",
input_references: [
{ type: "image_url", image_url: { url: "https://example.com/sign.jpg" } },
],
}),
});
console.log(await res.json());Can I mask the region to change?
| Field | Rule |
|---|---|
input_references | Source image; public HTTPS URL |
mask_url | Public HTTPS mask URL, documented for ChatGPT Image 2.5 edits |
aspect_ratio: "auto" | Matches the reference on edit calls |
| Unlisted field | 400 unsupported_parameter |
How do I keep the text exact?
Image models can misspell or restyle words. If the wording must be exact, generate or edit the picture first and set the text afterwards, as in n8n Edit Image text alignment. Models that handle text in generation are covered in text rendering models on Sume.
What should I check before using the result?
Read the new text letter by letter, then compare the rest of the image with the original for shifts in color or texture. Use n to request a few takes and pick the cleanest.
Sources
Related posts
More in Use cases
- AI music for one section of a video: no duration field
Sume's Music 1.0 rejects duration and duration_seconds. Ask for the length in the prompt, then loop and fade the track to fit one section on a Timeline.
- AI object remover from a photo: mask_url on GPT Image 2.5
To remove an object from a photo through the Sume API, send the photo and a public mask_url to POST /v1/images with a GPT Image 2.5 model, then prompt the fill.
- AI product color changer: colorways through an image edit API
To recolor a product photo, send it as an input_references image with a prompt naming the new color. Sume's edit calls take up to 10 images per request.
- Amazon AI-generated people tag: contains-synthetic-performer
Amazon's seller post says photorealistic AI people in listings, A+ and shoppable videos need the XMP keyword contains-synthetic-performer. Tag the final file.
Written by Sume