Translate a video to another language with AI voice, on time
A translated voice rarely lasts as long as the original. Use TTS sentence segments, the speed control and Timeline audio offsets to keep each line on its slot.

To translate a video into another language with an AI voice and keep the timing, treat each sentence as its own slot: read the source sentence times, speak the translation with sentence segmentation on, and compare each returned segment's length with its slot. Where a line runs long, shorten the text or raise generation_config.speed, which accepts 0.6 to 1.5. Sume does not stretch the voice to fit for you.
Which timing controls does Sume give me?
Four settings decide whether a dubbed line lands where the original did.
| Control | Value | Effect |
|---|---|---|
segmentation.mode | sentence, needs timestamps.words: true | Gapless segments[]; each end equals the next start |
boundary_lead_ms | 0 to 500, default 70 | Pause kept after a sentence's last word; the next segment absorbs it |
generation_config.speed | 0.6 to 1.5 | Speaking-rate multiplier for a line |
output_format.container | wav or raw | Each segment gets its own audio_url; mp3 returns timings only |
How do I tell whether a translated line fits?
Transcribe the original with segmentation.mode: "sentence"; each segment gives you start and end, so the slot is end minus start. Speak the translated sentences the same way and take the same subtraction on the TTS segments. If a line is only slightly longer than its slot, try a small speed increase and listen to it. If it is much longer, shorten the wording instead, since the speed control tops out at 1.5.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "Hoy presentamos el nuevo panel. Configurarlo lleva un minuto.",
"avatar_handle": "product_host",
"language": "es",
"generation_config": { "speed": 1.1 },
"output_format": { "container": "wav", "sample_rate": 44100 },
"timestamps": { "words": true },
"segmentation": { "mode": "sentence" }
}'How do I keep the joined audio in step with the picture?
Speak and re-time sentence by sentence, then join the files with POST /v1/timeline-1.0/audio using operation: "concat". The job returns segments[] with index, start and duration_seconds, which the docs describe as the concat offsets to re-base Timeline 1.0 video[].start against. Because the join is sample-domain, no silence is added at the seams, so any drift you see comes from the sentence lengths, not the join. Details are in the Timeline audio docs.
What is the order of work for a whole video?
Work in this order so that a fix never forces a full redo.
- Detach or inspect the source and transcribe it with sentence segments; save every slot as end minus start.
- Translate the sentences in your own code, one output line per source sentence, and check the count matches.
- Speak all lines with sentence segmentation on and a wav container, then list the lines whose segment is longer than the slot.
- Rewrite or speed up only those lines and speak them again; the files for the lines that fit stay as they are.
- Join the final files in order with a concat job and read the offsets from
segments[].
What if the voice and target language disagree?
Set language for every non-English transcript. The schema says it defaults to English when omitted, with Korean or Japanese inferred only from a Hangul- or kana-only transcript. Sume lists no per-language timing guarantees, so test one sentence per target language before you dub a whole video.
Sources
Related posts
More in Use cases
- Voice clone 10 seconds AI: what a short sample gets you on Sume
ElevenLabs says Instant Voice Clones can capture a voice from 10 seconds of audio. On Sume you clone in the app, from an upload or a 30-second recording.
- EU AI Act Article 50 AI video labeling requirements for creators
Article 50 applies from 2 August 2026: deployers must tell people about deepfakes and providers must add machine-readable marks. What that leaves to a creator.
- How to extend an AI video past 30 seconds
Sume has no extend parameter on video models. Chain clips: pull a last frame, use it as the next first_frame, then join with Timeline. Vendor limits too.
- Facebook Reels ad specs: 9:16, 1440x2560, H.264 and safe zones
Meta's ads guide for Facebook Reels lists 9:16, 1440x2560, MP4/MOV, 4 GB, H.264 and safe zones of 14% top and 35% bottom. What Sume can and cannot hit.
Written by Sume