Dub a video where two languages are spoken: keep the original lines

Sume's speech-to-text takes one language hint per call. For a mixed-language video, transcribe, translate only some lines, and splice the rest back in.

5 min readSume
All posts

To dub a video where two languages are spoken, do not replace the whole track. Transcribe it, decide line by line which sentences need a new voice, generate speech only for those, and splice the untouched stretches of the original audio back between them. Sume has the pieces for this, but none of them labels which language a sentence is in, so you make that call.

What follows comes from the Sume docs for video inspect, timeline audio and audio detach, read 2026-10-03.

What does Sume's transcription do with two languages?

Transcription takes one optional language_code, documented as a hint, such as en or ko. Leave it out and the language is auto-detected. With segmentation.mode: "sentence" the result also returns gapless sentence segments with timings, shaped like caption lines.

The docs describe a hint for the whole clip and do not describe a language label on each segment. So treat the transcript as text plus times, and decide yourself which lines are in which language. A second pass with a different hint can help on a file where one language dominates, but the docs promise nothing about it, so check the output by reading it.

How do you rebuild the track?

Timeline audio and the render accept up to 20 parts per call, and the join is sample-domain with no gap at the seams. Parts must share one channel layout, so keep the TTS output as wav at the same layout as the detached track, or the job fails with audio_parts_channel_mismatch.

  • Detach the audio once with audio detach, which returns a durable wav you can slice.
  • Transcribe with sentence segmentation and mark each sentence: translate it, or keep it.
  • For lines to translate, generate speech with TTS 1.0, setting the language and a voice for it.
  • Build one ordered list of parts: a TTS file for a translated line, and a slice of the detached original for a kept line, using source_in and duration.
  • Join the parts with Timeline audio concat, or put them straight on audio.parts[] of a Timeline render.

What are the limits of this approach?

Mixed-language dubbing limits on Sume, read 2026-10-03
LimitValueConsequence
Transcript language hintOne language_code per callNo per-sentence language labels documented
Transcript length hintUp to 600 secondsLonger files need a shorter source clip
Parts per join20 maximumA talky video needs grouped lines
Audio detachOutput up to 900 secondsDetach in ranges for long sources
TranslationNot in the public APIYou supply the translated lines
Background soundDetach returns the whole trackSpeech and music are not separated

What about the music and the room sound?

Detaching gives you the full audio track, with speech and anything else mixed together. Sume documents no speech and music separation, so when you replace a line with generated speech, the room sound that sat under the original line goes with it. A soundtrack bed added in the Timeline render, with duck_db to lower it under speech, is the documented way to bring music back.

For a clean result, the simplest case is a video where the spoken parts are in separate stretches, such as an interview with narration around it. Overlapping speakers cannot be split this way.

A worked example of the splice

Take a 90-second video where a host speaks English, and a guest answers in Korean in the middle. You detach the audio, transcribe with sentence segments, and find that sentences 1 to 6 are the host, 7 to 12 the guest, 13 to 16 the host again. To dub it into Spanish you translate all sixteen sentences, because the host and the guest both need to be understood by a Spanish audience.

To keep the guest's voice and add subtitles instead, you translate only the guest's lines into captions, and keep the original audio for that stretch as one part sliced with source_in and duration. The audio goes through unchanged and the captions carry the meaning. Which is better depends on whether hearing the guest matters more than a single language track; both are one join away.

In both cases the count matters. Sixteen sentences as sixteen parts fits within the 20-part limit, but a longer video does not, so group neighbouring sentences from the same speaker into one TTS job or one original slice before you build the list.

Checklist before you render

Read the transcript against the video once and mark the language of each line. Generate one translated line first and listen to it next to the original line before and after it, since a change in voice at the seam is what viewers notice. If you also need on-screen text in the new language, that is a separate caption job; see Muse Voice code-switching vs Sume STT auto-detect for the transcription side.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume