Localize one product ad into five languages: voice, captions
Keep the picture, swap the language: a per-language voiceover on sonic-3.6, burned captions with a language hint, and one Timeline render for each market.

To localize a product ad, keep one silent picture track and produce a new voiceover, captions and render for each market. On Sume that is POST /v1/tts-router/generate with a language per market, POST /v1/video-captions with a matching language hint, and POST /v1/timeline-1.0/render with the new voice as the audio spine.
Start from a clean master
Localization is cheaper when the picture has no baked-in speech and no burned text. Render the master cut without captions and with a neutral bed. Then every market reuses the same slots, and only the voice and the overlay change.
If your master already has a spoken track in one language, separate it first. The audio-detach and trim tools prepare material before assembly, which is the job split that Timeline documents.
One voice call for each language
The TTS Router takes a required model, a transcript and a voice selector. The language field comes from TTS 1.0. Sume also checks the language of the chosen voice against the request, and a known mismatch returns HTTP 409 tts_voice_language_mismatch. So choose a voice that speaks the target language, and do not send a Spanish script to an English-only voice.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ad42-es-voice" \
-d '{
"model": "sonic-3.6",
"transcript": "Tu sarten de hierro fundido, listo para toda la vida.",
"language": "es",
"voice": { "id": "YOUR_SPANISH_VOICE_UUID" }
}'Captions in the same language
For on-screen text, POST /v1/video-captions takes a public video_url, an optional language as a speech-to-text hint, an optional script_text, and a style. language only tells speech-to-text what to expect. It never picks the style or the font.
For Korean, the docs add a style named korean-ad, which is CapCut-style Hangul karaoke, one short phrase at a time. If you pass no style, Korean text resolves to black-outline, because the Latin default would render Hangul as tofu.
Assemble and name your keys
Render each market with the same video[] slots and its own audio.url and audio.duration_seconds. Name the idempotency keys by ad and language, such as ad42-es-timeline, so you can find a market's job later.
Check one thing before you ship: a translated line is usually longer or shorter than the original. Because the voice sets the audio spine, the render length follows the new voice, and the slots may need new durations. Use POST /v1/timeline-1.0/plan to see duration_seconds and billable_minutes before you pay.
Sources
Related posts
More in Use cases
- Localize a YouTube thumbnail into 3 languages for $0.11 on Sume
One finished thumbnail, three Ideogram 4.5 edits through POST /v1/images: Spanish, Portuguese, German at low quality for about $0.11, with the prompt and code.
- Logo sketch to a transparent PNG with GPT Image 2.5
Turn a hand-drawn logo sketch into a transparent PNG: send the sketch as a reference, set background transparent and output_format png, then verify alpha.
- Voice drift in long narration: Nova 2 Sonic's 52% vs Sume chunks
Speaker drift is what splits a long voiceover into jobs. Amazon claims -52% on an internal set; here is how to measure drift across Sume TTS chunks yourself.
- Luggage holiday travel ad: fit a 16:9 clip to 9:16 with blur for $0.11
Reuse a 16:9 luggage ad as a 9:16 Reel: detach the audio for $0.01, then one Timeline render with fit blur for $0.10, instead of $1.00 to regenerate.
Written by Sume