One video in three languages: per-language voiceover tracks on Sume

Render scene lines in each language with TTS, join each language into one track with timeline audio concat, then re-base scene starts from segments[].

6 min readSume
All posts

For one product video in English, Spanish and German, render each scene line as its own TTS job per language, join each language's lines into a single track with timeline audio concat, and render three Timeline 1.0 videos from the same footage. The concat result returns segments[] with each line's offset, which is what you re-base scene start times against, because German rarely takes the same time as English. Cost is per job, so the plan is easy to price.

Why per-line jobs instead of one long script

A single long transcript gives you one audio file and no scene boundaries. Per-line jobs give you a file per scene that you can re-record alone. Concat then produces a gapless track in the sample domain, with no re-synthesis and no silence at the joins, according to the timeline audio docs. Parts must share one channel layout; a mismatch is refused with audio_parts_channel_mismatch, so request the same output format for every line.

The counts

Take 6 scenes and 3 languages.

Job count for 6 scenes in 3 languages (Sume docs, read 2026-10-04)
StepJobsNote
TTS per scene per language18One job per line; pick the voice and language per job (TTS parameters are not covered in the docs read here)
Timeline audio concat per language3parts[] up to 20; $0.01 flat per job
Timeline render per language3$0.10 per output minute, rounded up

The concat call

Import or reuse the TTS artifacts so each URL is on media.sume.com, then send the parts in scene order. Use WAV when the file will be joined again.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: es-track-001" \
  -d '{
    "operation": "concat",
    "parts": [
      { "url": "https://media.sume.com/artifacts/artf_demo/es-scene1.wav" },
      { "url": "https://media.sume.com/artifacts/artf_demo/es-scene2.wav" }
    ]
  }'

Re-basing the scene starts

The result has one audio_url, a duration_seconds, and segments[] with index, start and duration_seconds. In Timeline 1.0, video[0].start must be 0 and later starts must increase, and each slot's coverage may trail the spine by at most 0.5 seconds. So set each video[i].start to the matching segment start, and give each slot a duration that covers its line. A German line that runs 1.4 seconds longer pushes every later scene; that is exactly what the segments array tells you.

Microsoft's MAI-Voice-2.1 page lists 23 supported languages (read 2026-10-04) but does not say whether one voice carries a native accent across all of them, so audition a voice per language, as in the 23-language test matrix.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume