Dub a video where two languages are spoken: keep the original lines
Sume's speech-to-text takes one language hint per call. For a mixed-language video, transcribe, translate only some lines, and splice the rest back in.

To dub a video where two languages are spoken, do not replace the whole track. Transcribe it, decide line by line which sentences need a new voice, generate speech only for those, and splice the untouched stretches of the original audio back between them. Sume has the pieces for this, but none of them labels which language a sentence is in, so you make that call.
What follows comes from the Sume docs for video inspect, timeline audio and audio detach, read 2026-10-03.
What does Sume's transcription do with two languages?
Transcription takes one optional language_code, documented as a hint, such as en or ko. Leave it out and the language is auto-detected. With segmentation.mode: "sentence" the result also returns gapless sentence segments with timings, shaped like caption lines.
The docs describe a hint for the whole clip and do not describe a language label on each segment. So treat the transcript as text plus times, and decide yourself which lines are in which language. A second pass with a different hint can help on a file where one language dominates, but the docs promise nothing about it, so check the output by reading it.
How do you rebuild the track?
Timeline audio and the render accept up to 20 parts per call, and the join is sample-domain with no gap at the seams. Parts must share one channel layout, so keep the TTS output as wav at the same layout as the detached track, or the job fails with audio_parts_channel_mismatch.
- Detach the audio once with audio detach, which returns a durable wav you can slice.
- Transcribe with sentence segmentation and mark each sentence: translate it, or keep it.
- For lines to translate, generate speech with TTS 1.0, setting the
languageand a voice for it. - Build one ordered list of parts: a TTS file for a translated line, and a slice of the detached original for a kept line, using
source_inandduration. - Join the parts with Timeline audio concat, or put them straight on
audio.parts[]of a Timeline render.
What are the limits of this approach?
| Limit | Value | Consequence |
|---|---|---|
| Transcript language hint | One language_code per call | No per-sentence language labels documented |
| Transcript length hint | Up to 600 seconds | Longer files need a shorter source clip |
| Parts per join | 20 maximum | A talky video needs grouped lines |
| Audio detach | Output up to 900 seconds | Detach in ranges for long sources |
| Translation | Not in the public API | You supply the translated lines |
| Background sound | Detach returns the whole track | Speech and music are not separated |
What about the music and the room sound?
Detaching gives you the full audio track, with speech and anything else mixed together. Sume documents no speech and music separation, so when you replace a line with generated speech, the room sound that sat under the original line goes with it. A soundtrack bed added in the Timeline render, with duck_db to lower it under speech, is the documented way to bring music back.
For a clean result, the simplest case is a video where the spoken parts are in separate stretches, such as an interview with narration around it. Overlapping speakers cannot be split this way.
A worked example of the splice
Take a 90-second video where a host speaks English, and a guest answers in Korean in the middle. You detach the audio, transcribe with sentence segments, and find that sentences 1 to 6 are the host, 7 to 12 the guest, 13 to 16 the host again. To dub it into Spanish you translate all sixteen sentences, because the host and the guest both need to be understood by a Spanish audience.
To keep the guest's voice and add subtitles instead, you translate only the guest's lines into captions, and keep the original audio for that stretch as one part sliced with source_in and duration. The audio goes through unchanged and the captions carry the meaning. Which is better depends on whether hearing the guest matters more than a single language track; both are one join away.
In both cases the count matters. Sixteen sentences as sixteen parts fits within the 20-part limit, but a longer video does not, so group neighbouring sentences from the same speaker into one TTS job or one original slice before you build the list.
Checklist before you render
Read the transcript against the video once and mark the language of each line. Generate one translated line first and listen to it next to the original line before and after it, since a change in voice at the seam is what viewers notice. If you also need on-screen text in the new language, that is a separate caption job; see Muse Voice code-switching vs Sume STT auto-detect for the transcription side.
Sources
Related posts
More in Use cases
- Edit one region of an AI image and keep every other pixel
A mask edit is guidance, not a guarantee. Paste the original back outside the mask with Pillow so untouched pixels stay identical. Python with Sume's mask_url.
- Face swap on someone else's Short: does YouTube call it original?
YouTube said Oct 1 that Shorts re-uploading others' videos without significant changes will get less reach. Why a face swap on borrowed footage is a risk.
- GPT Image 2.5 icon set: one anchor icon, then reference it
Make a consistent icon set with GPT Image 2.5 on Sume: approve one anchor icon, pass it as input_references, change only the subject, and keep a style block.
- GPT Image 2.5 pixel art sprite sheet: prompt a grid, slice it
Ask GPT Image 2.5 for a 4x2 sprite sheet at 2048x1024 on Sume, then slice frames and snap to a true pixel grid with Pillow. Prompt and code included.
Written by Sume