Multi-speaker dub: 32 speakers on ElevenLabs, a voice per line on Sume
ElevenLabs dubbing handles up to 32 speakers per file. On Sume you give each speaker a voice, make TTS jobs per line, and join up to 20 parts per concat.

ElevenLabs dubbing v2 handles up to 32 speakers in one file automatically. Sume documents no one-call dubbing endpoint, so a multi-speaker dub there is manual: assign each speaker a voice, generate each translated line with its own TTS job, and join the lines in order with timeline audio concat, which takes 1 to 20 parts per call.
It is more work than a one-click dub, and in return you control every line.
What the vendor does
The ElevenLabs dubbing doc lists up to 32 speakers per file, 90+ languages in v2, and file limits of 1 GB and 180 minutes in the app or 3 GB per file via API. It notes that editing the transcript and regenerating through the API is Enterprise only, which matters if you want to fix one speaker's line.
The Sume build
Start with a script that marks which speaker says each line and when. Use sentence segmentation or your own timings to know each slot. Generate one TTS job per line with the voice for that speaker and the target language, then join them. The timeline audio docs say concat takes 1 to 20 parts, preserves them in the sample domain with no re-synthesis, and returns segments[] with the offsets of each part so that you can line video cuts up with the new audio.
- One voice id per speaker, kept constant across the whole video.
- One language per job; set the language code explicitly.
- Match every part's channel layout to avoid audio_parts_channel_mismatch.
- Past 20 lines, concat in groups and concat the group outputs, since results are hosted files.
Timing and fit
A translation is rarely the same length as the original. Use the TTS speed field (0.6 to 1.5) to nudge a line to fit its slot, and rewrite the line if it still overruns. Each concat part takes a url with optional source_in and duration, so you can trim a take before it is joined.
| Step | ElevenLabs dubbing | Sume |
|---|---|---|
| Detect speakers | Automatic, up to 32 | You mark speakers in the script |
| Translate | Included | You or an agent write the translation |
| Voices | Per speaker, automatic | One voice id per speaker, chosen by you |
| Fix one line | Enterprise API feature | Regenerate one TTS job and re-run concat |
| Join | Included | Timeline audio concat, 1 to 20 parts |
When this is worth it
For short videos with two or three speakers, or when you must approve each translated line, the manual path is practical and cheap to correct. For a long panel with many voices, use a dedicated dubbing product and bring the finished audio into Sume only for captions and final composition.
Sources
Related posts
More in Use cases
- Music bed longer than your Short: YouTube Create caps it, Sume cuts it
YouTube Create says audio cannot exceed the video's length. A Sume Timeline render ends at audio.duration_seconds, with a soundtrack fade up to 10 s.
- Music for an avatar video: pass the preview still as image_url
Approve an avatar preview, then send its preview_image_url to the Music Router as image_url so the track matches the frame. Steps, limits and the $0.125 cost.
- Nano Banana 2 catalog photos: draft at 0.5K, finish at 2K
Cut catalog image cost by drafting at 0.5K and rendering only approved shots at 2K with Nano Banana 2; vendor rates and the Sume resolution tiers.
- Narrate a 3-minute Short: Sume TTS, word timings and 60 s caption cuts
YouTube allows three-minute Shorts. A plan with Sume: one TTS request under 20,000 characters, word timings, and caption jobs in pieces of 60 seconds or less.
Written by Sume