Localize one product ad into five languages: voice, captions

Keep the picture, swap the language: a per-language voiceover on sonic-3.6, burned captions with a language hint, and one Timeline render for each market.

5 min readSume
All posts

To localize a product ad, keep one silent picture track and produce a new voiceover, captions and render for each market. On Sume that is POST /v1/tts-router/generate with a language per market, POST /v1/video-captions with a matching language hint, and POST /v1/timeline-1.0/render with the new voice as the audio spine.

Start from a clean master

Localization is cheaper when the picture has no baked-in speech and no burned text. Render the master cut without captions and with a neutral bed. Then every market reuses the same slots, and only the voice and the overlay change.

If your master already has a spoken track in one language, separate it first. The audio-detach and trim tools prepare material before assembly, which is the job split that Timeline documents.

One voice call for each language

The TTS Router takes a required model, a transcript and a voice selector. The language field comes from TTS 1.0. Sume also checks the language of the chosen voice against the request, and a known mismatch returns HTTP 409 tts_voice_language_mismatch. So choose a voice that speaks the target language, and do not send a Spanish script to an English-only voice.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ad42-es-voice" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "Tu sarten de hierro fundido, listo para toda la vida.",
    "language": "es",
    "voice": { "id": "YOUR_SPANISH_VOICE_UUID" }
  }'

Captions in the same language

For on-screen text, POST /v1/video-captions takes a public video_url, an optional language as a speech-to-text hint, an optional script_text, and a style. language only tells speech-to-text what to expect. It never picks the style or the font.

For Korean, the docs add a style named korean-ad, which is CapCut-style Hangul karaoke, one short phrase at a time. If you pass no style, Korean text resolves to black-outline, because the Latin default would render Hangul as tofu.

Assemble and name your keys

Render each market with the same video[] slots and its own audio.url and audio.duration_seconds. Name the idempotency keys by ad and language, such as ad42-es-timeline, so you can find a market's job later.

Check one thing before you ship: a translated line is usually longer or shorter than the original. Because the voice sets the audio spine, the render length follows the new voice, and the slots may need new durations. Use POST /v1/timeline-1.0/plan to see duration_seconds and billable_minutes before you pay.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume