Translate a video to another language with AI voice, on time

A translated voice rarely lasts as long as the original. Use TTS sentence segments, the speed control and Timeline audio offsets to keep each line on its slot.

4 min readSume
All posts

To translate a video into another language with an AI voice and keep the timing, treat each sentence as its own slot: read the source sentence times, speak the translation with sentence segmentation on, and compare each returned segment's length with its slot. Where a line runs long, shorten the text or raise generation_config.speed, which accepts 0.6 to 1.5. Sume does not stretch the voice to fit for you.

Which timing controls does Sume give me?

Four settings decide whether a dubbed line lands where the original did.

Timing controls in TTS 1.0 and Timeline audio, read 2026-09-29.
ControlValueEffect
segmentation.modesentence, needs timestamps.words: trueGapless segments[]; each end equals the next start
boundary_lead_ms0 to 500, default 70Pause kept after a sentence's last word; the next segment absorbs it
generation_config.speed0.6 to 1.5Speaking-rate multiplier for a line
output_format.containerwav or rawEach segment gets its own audio_url; mp3 returns timings only

How do I tell whether a translated line fits?

Transcribe the original with segmentation.mode: "sentence"; each segment gives you start and end, so the slot is end minus start. Speak the translated sentences the same way and take the same subtraction on the TTS segments. If a line is only slightly longer than its slot, try a small speed increase and listen to it. If it is much longer, shorten the wording instead, since the speed control tops out at 1.5.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "transcript": "Hoy presentamos el nuevo panel. Configurarlo lleva un minuto.",
    "avatar_handle": "product_host",
    "language": "es",
    "generation_config": { "speed": 1.1 },
    "output_format": { "container": "wav", "sample_rate": 44100 },
    "timestamps": { "words": true },
    "segmentation": { "mode": "sentence" }
  }'

How do I keep the joined audio in step with the picture?

Speak and re-time sentence by sentence, then join the files with POST /v1/timeline-1.0/audio using operation: "concat". The job returns segments[] with index, start and duration_seconds, which the docs describe as the concat offsets to re-base Timeline 1.0 video[].start against. Because the join is sample-domain, no silence is added at the seams, so any drift you see comes from the sentence lengths, not the join. Details are in the Timeline audio docs.

What is the order of work for a whole video?

Work in this order so that a fix never forces a full redo.

  • Detach or inspect the source and transcribe it with sentence segments; save every slot as end minus start.
  • Translate the sentences in your own code, one output line per source sentence, and check the count matches.
  • Speak all lines with sentence segmentation on and a wav container, then list the lines whose segment is longer than the slot.
  • Rewrite or speed up only those lines and speak them again; the files for the lines that fit stay as they are.
  • Join the final files in order with a concat job and read the offsets from segments[].

What if the voice and target language disagree?

Set language for every non-English transcript. The schema says it defaults to English when omitted, with Korean or Japanese inferred only from a Hangul- or kana-only transcript. Sume lists no per-language timing guarantees, so test one sentence per target language before you dub a whole video.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume