AI change text in image: an edit call with a reference

To change wording inside a picture, send it as an input_references image with a prompt giving the new text. Sume lists edit-capable models.

4 min readSume
All posts

Send the picture as an input_references image to POST /v1/images with a prompt that quotes the old wording and the new wording, and use aspect_ratio: "auto" to keep its shape. The example uses google/nano-banana-2, which is edit-capable and lists auto; Sume also lists Ideogram V3 as edit-capable, but its ratio list has no auto. Results can change more than the text, so check them; for exact wording, overlay the text in your own tool.

Request rules are from Image generation and editing, read 2026-09-30. Ideogram's 4.5 page lists text modification among its edit uses; Sume's catalog lists V3, not 4.5.

What does the call look like?

Name the text to replace and the text to use, and tell the model to leave the rest alone.

const res = await fetch("https://api.sume.com/v1/images", {
  method: "POST",
  headers: {
    Authorization: "Bearer " + process.env.SUME_API_KEY,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "google/nano-banana-2",
    prompt: 'Change the sign text from "OPEN" to "SALE". Keep everything else unchanged.',
    aspect_ratio: "auto",
    input_references: [
      { type: "image_url", image_url: { url: "https://example.com/sign.jpg" } },
    ],
  }),
});
console.log(await res.json());

Can I mask the region to change?

Edit fields in the Sume docs, read 2026-09-30.
FieldRule
input_referencesSource image; public HTTPS URL
mask_urlPublic HTTPS mask URL, documented for ChatGPT Image 2.5 edits
aspect_ratio: "auto"Matches the reference on edit calls
Unlisted field400 unsupported_parameter

How do I keep the text exact?

Image models can misspell or restyle words. If the wording must be exact, generate or edit the picture first and set the text afterwards, as in n8n Edit Image text alignment. Models that handle text in generation are covered in text rendering models on Sume.

What should I check before using the result?

Read the new text letter by letter, then compare the rest of the image with the original for shifts in color or texture. Use n to request a few takes and pick the cleanest.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume