Multi-speaker dub: 32 speakers on ElevenLabs, a voice per line on Sume

ElevenLabs dubbing handles up to 32 speakers per file. On Sume you give each speaker a voice, make TTS jobs per line, and join up to 20 parts per concat.

4 min readSume
All posts

ElevenLabs dubbing v2 handles up to 32 speakers in one file automatically. Sume documents no one-call dubbing endpoint, so a multi-speaker dub there is manual: assign each speaker a voice, generate each translated line with its own TTS job, and join the lines in order with timeline audio concat, which takes 1 to 20 parts per call.

It is more work than a one-click dub, and in return you control every line.

What the vendor does

The ElevenLabs dubbing doc lists up to 32 speakers per file, 90+ languages in v2, and file limits of 1 GB and 180 minutes in the app or 3 GB per file via API. It notes that editing the transcript and regenerating through the API is Enterprise only, which matters if you want to fix one speaker's line.

The Sume build

Start with a script that marks which speaker says each line and when. Use sentence segmentation or your own timings to know each slot. Generate one TTS job per line with the voice for that speaker and the target language, then join them. The timeline audio docs say concat takes 1 to 20 parts, preserves them in the sample domain with no re-synthesis, and returns segments[] with the offsets of each part so that you can line video cuts up with the new audio.

  • One voice id per speaker, kept constant across the whole video.
  • One language per job; set the language code explicitly.
  • Match every part's channel layout to avoid audio_parts_channel_mismatch.
  • Past 20 lines, concat in groups and concat the group outputs, since results are hosted files.

Timing and fit

A translation is rarely the same length as the original. Use the TTS speed field (0.6 to 1.5) to nudge a line to fit its slot, and rewrite the line if it still overruns. Each concat part takes a url with optional source_in and duration, so you can trim a take before it is joined.

Dub responsibilities: vendor automatic dubbing versus a Sume build (read 2026-10-03)
StepElevenLabs dubbingSume
Detect speakersAutomatic, up to 32You mark speakers in the script
TranslateIncludedYou or an agent write the translation
VoicesPer speaker, automaticOne voice id per speaker, chosen by you
Fix one lineEnterprise API featureRegenerate one TTS job and re-run concat
JoinIncludedTimeline audio concat, 1 to 20 parts

When this is worth it

For short videos with two or three speakers, or when you must approve each translated line, the manual path is practical and cheap to correct. For a long panel with many voices, use a dedicated dubbing product and bring the finished audio into Sume only for captions and final composition.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume