Localize app store screenshot captions: 6 locales x 5 shots, $1.13
Translate the caption on 5 marketing screenshots into 6 locales with Ideogram 4.5 edits on Sume: 30 calls, $1.125 at low, and why real UI gets re-captured.

To localize app store screenshot captions with AI, edit only the marketing line above or below the device frame: one Ideogram 4.5 call per screenshot per locale through POST /v1/images, with the finished screenshot as the first input_references image. Five screenshots in six locales is 30 edits, which cost $1.125 at low quality and $2.25 at medium on Sume.
One rule decides whether this is a good idea: the caption is translated by the model, but the app UI inside the device frame is not yours to retype. If the phone shows a screen with English labels, that is a picture of your product, and the honest localized version is a new screenshot from the localized build. Use edits for the words around the phone, and capture the real UI separately.
What does one call send?
Quote the old caption and the new one, name the area, and fence off the device: "Replace only the caption text above the phone, from 'Plan your week' to 'Planifica tu semana'. Keep the font, color and position. Do not change the phone or anything on its screen." Send no aspect_ratio, and the edit keeps the shape of the source, per the Image API docs, so your store-size master stays the right size.
Ideogram's launch post calls 4.5 an edit model that keeps multi-turn work free of artifacts (Ideogram on X, read 2026-10-05), which is the property you want when the phone screen must not move. Check it yourself: put the source and result on top of each other and look at the device frame first.
What does 30 edits cost?
Ideogram 4.5 is priced by quality and not by size: $0.03, $0.06 or $0.22 list, times 1.25 at Sume. The matrix below uses five screenshots.
| Quality | Per edit | 6 locales (30 edits) | 12 locales (60 edits) |
|---|---|---|---|
| low | $0.0375 | $1.125 | $2.25 |
| medium | $0.075 | $2.25 | $4.50 |
| high | $0.275 | $8.25 | $16.50 |
How do I run 30 edits in a sane order?
Run one locale first, all five screenshots, and read them before you queue the other five locales. Fixing a prompt after five edits costs $0.19 at low; fixing it after 30 costs the full rerun. Sume accepts valid paid jobs as queued within your plan's capacity (Free accepts 6 at once, Pro 24, Startup 48, per Generation admission), so on a Free plan, a locale of five fits in one wave.
Use keys like shot2-es-v1. A key that repeats with the same payload returns the original job instead of billing twice, and a new prompt needs a new key (shot2-es-v2).
What fails first in captions?
Run each result past a native reader, not a spell checker. These are the failures to look for:
- Longer text. German and Finnish captions often run well past the English length. The model may shrink the font or wrap the line; the master layout is where you fix it, by shortening the translation.
- Right-to-left scripts. If you target Arabic or Hebrew, test a single shot first and read the direction and letter joins by eye.
- Non-Latin glyphs. Check Japanese, Thai and Hindi at
mediumbefore you commit tolowfor the set. - File format. Flatten the outputs to PNG without alpha before upload; see flatten an RGBA PNG before upload.
Total spend for a trial that tells you if this works: one locale of five at low is $0.1875, plus a rerun of the two worst at medium for $0.15. That is $0.34 to know whether to run the other 25.
A small test before the full run
Before you scale this up, run it on two or three real files first and write down what you saw. A small test at low quality costs cents on Sume ($0.0375 per Ideogram 4.5 edit), and it tells you whether your prompt, your source files and your review step are ready.
Keep the originals untouched, name every output after its source and its prompt, and store the job id with each result. If a result is wrong later, you can find the exact request, fix the prompt and re-run only that item with a new Idempotency-Key.
Sources
Related posts
More in Use cases
- Localize one 30-second ad into six languages: Wan 3.0 cost on Sume
Six language versions of a 30 second ad on Wan 3.0 cost $22.50 at 720p or $45.00 at 1080p on Sume. Draft all six at 480p for $11.25 first.
- Localize one ad into 14 languages: TTS, captions and render for $4.62
One 450-character ad in 14 shared languages costs $0.42 TTS, $2.80 captions and $1.40 render on Sume: $4.62 total. The loop uses one job per language.
- Localize a YouTube thumbnail into 3 languages for $0.11 on Sume
One finished thumbnail, three Ideogram 4.5 edits through POST /v1/images: Spanish, Portuguese, German at low quality for about $0.11, with the prompt and code.
- Logo sketch to a transparent PNG with GPT Image 2.5
Turn a hand-drawn logo sketch into a transparent PNG: send the sketch as a reference, set background transparent and output_format png, then verify alpha.
Written by Sume