Translate text inside an image: one Nano Banana 2 edit per language
Google says Nano Banana 2 can translate and localize text within an image. Loop one edit per language through Sume POST /v1/images and keep the layout fixed.

To localize an image that has text in it, send the original as an input_references URL to google/nano-banana-2 with one instruction per language, and set aspect_ratio to auto so the layout is not recropped. Google's launch post says Nano Banana 2 can 'translate and localize text within an image' (Google, read 2026-10-01). Sume exposes the model on POST /v1/images, so a loop of N languages is N edit calls.
This is an edit, not a redraw: you are asking the model to keep the artwork and swap the words. Check every output, because short headlines survive better than dense paragraphs.
What do I send?
The fields below come from the Image API docs.
| Field | Value | Why |
|---|---|---|
model | google/nano-banana-2 | The model Google describes for text localization |
input_references | One public HTTPS image URL | The source artwork |
aspect_ratio | auto | Docs advise auto on edits to match the input |
n | 1 | One result per language, easy to label |
What does the loop look like?
Keep the instruction short and name the exact strings you want, so the model does not invent copy. Each call can return a hosted URL in data[0].url.
import os
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
SRC = os.environ["SOURCE_IMAGE_URL"]
LANGS = ["Spanish", "German", "Japanese"]
for lang in LANGS:
r = requests.post("https://api.sume.com/v1/images", headers=H, timeout=60,
json={"model": "google/nano-banana-2", "aspect_ratio": "auto",
"prompt": f"Translate all text in this image to {lang}. "
"Keep the artwork, fonts and layout unchanged.",
"input_references": [{"type": "image_url",
"image_url": {"url": SRC}}]})
print(lang, r.status_code, r.json().get("data", [{}])[0].get("url"))How do I check the result?
Have a fluent reader look at each file. Models can swap characters in scripts with many glyphs, and Japanese or Arabic need the most scrutiny. If a language fails, retry that language alone rather than the batch. The cost of each call is in usage.cost; add them up before committing to a large language list.
Limits
Google's claim is from its own post and I did not test accuracy per language. A 202 job envelope means the call outlasted the 30-second sync wait; poll the job. The model may reflow text that is longer in the target language, so leave room in the original design.
Sources
Related posts
More in Use cases
- Pin a product card on a live-selling clip with one overlay call
POST /v1/timeline-1.0/compose with operation overlay pins a product still at bottom, center or top of a live-selling clip for $0.02. Layout keys and limits.
- Podcast audiogram API: audio plus a cover still in one render
Make a podcast audiogram video from an audio clip and a cover image with Timeline 1.0: spine, still slot, square output and the cost per minute.
- Turn a product changelog into a 30-second update video
Convert release notes into a short update video: pick three changes, write a 70-word script, voice it with TTS, add real screenshots and render with Timeline.
- Launch teaser from product shots: Gemini Omni Flash reference images
Feed up to 10 product references to gemini-omni-flash-1.1 on Sume, address them as IMAGE_REF_0 in the prompt, and get a 3-10 second teaser with native audio.
Written by Sume