One video in three languages: per-language voiceover tracks on Sume
Render scene lines in each language with TTS, join each language into one track with timeline audio concat, then re-base scene starts from segments[].

For one product video in English, Spanish and German, render each scene line as its own TTS job per language, join each language's lines into a single track with timeline audio concat, and render three Timeline 1.0 videos from the same footage. The concat result returns segments[] with each line's offset, which is what you re-base scene start times against, because German rarely takes the same time as English. Cost is per job, so the plan is easy to price.
Why per-line jobs instead of one long script
A single long transcript gives you one audio file and no scene boundaries. Per-line jobs give you a file per scene that you can re-record alone. Concat then produces a gapless track in the sample domain, with no re-synthesis and no silence at the joins, according to the timeline audio docs. Parts must share one channel layout; a mismatch is refused with audio_parts_channel_mismatch, so request the same output format for every line.
The counts
Take 6 scenes and 3 languages.
| Step | Jobs | Note |
|---|---|---|
| TTS per scene per language | 18 | One job per line; pick the voice and language per job (TTS parameters are not covered in the docs read here) |
| Timeline audio concat per language | 3 | parts[] up to 20; $0.01 flat per job |
| Timeline render per language | 3 | $0.10 per output minute, rounded up |
The concat call
Import or reuse the TTS artifacts so each URL is on media.sume.com, then send the parts in scene order. Use WAV when the file will be joined again.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: es-track-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/es-scene1.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/es-scene2.wav" }
]
}'Re-basing the scene starts
The result has one audio_url, a duration_seconds, and segments[] with index, start and duration_seconds. In Timeline 1.0, video[0].start must be 0 and later starts must increase, and each slot's coverage may trail the spine by at most 0.5 seconds. So set each video[i].start to the matching segment start, and give each slot a duration that covers its line. A German line that runs 1.4 seconds longer pushes every later scene; that is exactly what the segments array tells you.
Microsoft's MAI-Voice-2.1 page lists 23 supported languages (read 2026-10-04) but does not say whether one voice carries a native accent across all of them, so audition a voice per language, as in the 23-language test matrix.
Sources
Related posts
More in Use cases
- Online course trailer in 30 seconds with Seedance 2.5: beats and cost
A 30-second course trailer in one Seedance 2.5 request: beats for hook, promise and call to action, authored captions on top, about $17.33 at 720p on Sume.
- Out-of-office video message with an AI avatar: 15 seconds on Sume
Make a 15-second out-of-office video from a script: return date, contact, captions on. $2.76 on Standard plus a $0.20 caption add-on; Plus and Max too.
- Parent-teacher conference reminder in English and Spanish for $0.63
A conference reminder from one slide image in two languages: TTS, a Timeline render and captions per version. $0.6333 for both on Sume, from docs.
- Peacock vertical video ads in 2026: what to prepare before specs land
NBCUniversal says Peacock's vertical video format opens to advertisers in 2026 but lists no specs. Prepare a 9:16 master with Sume's timeline and wait.
Written by Sume