Localize a YouTube thumbnail into 3 languages for $0.11 on Sume
One finished thumbnail, three Ideogram 4.5 edits through POST /v1/images: Spanish, Portuguese, German at low quality for about $0.11, with the prompt and code.
To localize a YouTube thumbnail with AI, send the finished thumbnail as the first input_references image to POST /v1/images with model: "ideogram/ideogram-v4.5" and a prompt that quotes the old headline and the new one. Three languages cost about $0.11 at low quality, because Sume bills Ideogram 4.5 at its list price times 1.25: $0.0375 per edit at low, $0.075 at medium, $0.275 at high.
Ideogram described 4.5 on 2026-09-30 as an edit model that "eliminates artifact buildup" (Ideogram on X, read 2026-10-05). That claim is the reason to try it for this job: you want the face, the arrow and the background to stay put while only the words change. Treat it as a vendor claim, and run your own thumbnail through it before you ship a channel.
What does one localized thumbnail request look like?
The first reference is the image Ideogram edits. Sume's Image API docs say an edit without aspect_ratio keeps the shape of the source image, so a 16:9 thumbnail comes back 16:9 and you do not have to name a ratio. Quality is low, medium or high, and medium is the default if you leave it out.
Quote both strings in the prompt and say what must not change. Add a line that gives the target language, because the model has no other signal about which accent marks you expect.
import json, os, urllib.request
SRC = "https://example.com/thumb-en.png" # public https URL
LANGS = {"es": "PIERDE 5 KILOS", "pt": "PERCA 5 QUILOS", "de": "VERLIER 5 KILO"}
def edit(lang, headline):
body = {
"model": "ideogram/ideogram-v4.5",
"quality": "low",
"prompt": f'Replace the headline "LOSE 5 KILOS" with "{headline}". '
"Same font, colour, size and position. Change nothing else.",
"input_references": [{"type": "image_url", "image_url": {"url": SRC}}],
}
req = urllib.request.Request(
"https://api.sume.com/v1/images", json.dumps(body).encode(),
{"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": f"thumb-001-{lang}"})
with urllib.request.urlopen(req, timeout=60) as r:
return r.status, json.load(r)
for lang, headline in LANGS.items():
status, body = edit(lang, headline)
print(lang, status, body["data"][0]["url"] if status == 200 else body)What does it cost per video and per channel?
Each language is one edit and one image, and Ideogram 4.5's price does not change with size, so a 1K and a 2K thumbnail cost the same. The table prices one video in three languages and a ten-video batch.
| Quality | Per edit | 1 video, 3 languages | 10 videos, 3 languages (30 edits) |
|---|---|---|---|
| low | $0.0375 | $0.1125 | $1.125 |
| medium | $0.075 | $0.225 | $2.25 |
| high | $0.275 | $0.825 | $8.25 |
Which quality should a thumbnail edit use?
Start at low for every language, open the three results side by side, and re-run only the ones that fail. A thumbnail has a short headline, so low is often enough; the failure you are looking for is a mangled letter or a changed font, not a soft texture. Re-running one failed German line at medium costs $0.075, which is still cheaper than one high edit at $0.275.
A launch-partner page for Ideogram 4.5 (Morphic, a partner and not the vendor, read 2026-10-05) says the model edits stylized text in place and translates it, and copies unchanged pixels from the source. That matches the job, but it is a partner description. The Sume docs make no promise about text quality, so your check is the only proof.
What breaks in translation and what to check
Check these four things on every result before upload:
- Length. German and Portuguese run longer than English, so a headline that fit in 14 characters may not fit in 20. Shorten the translated line before you send it, not after.
- Accents. Look at tildes, cedillas and umlauts letter by letter; they are the first thing to fail.
- Everything else. Compare the face, logo and background to the source. Anything that moved is a failed edit, whatever the headline says.
- Status code.
POST /v1/imageswaits up to 30 seconds and returns200. Slow runs fall back to202with a job envelope, so branch on the code and pollstatus_urlas described in Jobs and results.
How do I keep the three jobs from billing twice?
Send a stable Idempotency-Key per language, as the code does (thumb-001-es). If a retry reuses the key with the same payload, you get the original job back and no second charge. Change the prompt and you must change the key. The YouTube localized thumbnails checklist covers the upload side.
A small test before the full run
Before you scale this up, run it on two or three real files first and write down what you saw. A small test at low quality costs cents on Sume ($0.0375 per Ideogram 4.5 edit), and it tells you whether your prompt, your source files and your review step are ready.
Keep the originals untouched, name every output after its source and its prompt, and store the job id with each result. If a result is wrong later, you can find the exact request, fix the prompt and re-run only that item with a new Idempotency-Key.
Sources
Related posts
More in Use cases
- Logo sketch to a transparent PNG with GPT Image 2.5
Turn a hand-drawn logo sketch into a transparent PNG: send the sketch as a reference, set background transparent and output_format png, then verify alpha.
- Voice drift in long narration: Nova 2 Sonic's 52% vs Sume chunks
Speaker drift is what splits a long voiceover into jobs. Amazon claims -52% on an internal set; here is how to measure drift across Sume TTS chunks yourself.
- Luggage holiday travel ad: fit a 16:9 clip to 9:16 with blur for $0.11
Reuse a 16:9 luggage ad as a 9:16 Reel: detach the audio for $0.01, then one Timeline render with fit blur for $0.10, instead of $1.00 to regenerate.
- Live transcription for Korean meetings: MAI-Transcribe-2-Streaming
MAI-Transcribe-2-Streaming claims 60 languages with auto detection and about 100ms to first text. Verify Korean yourself. Sume transcribes files, not live.
Written by Sume