ElevenLabs dubbing 32 speakers per file vs Sume one voice per TTS job
ElevenLabs dubbing handles up to 32 speakers in a file. Sume composes dubbing from STT and TTS, one voice per job. What that means for multi-speaker video.

What the vendor page says
The ElevenLabs dubbing page says a file can contain up to 32 speakers and that dubbing covers 90+ languages. It also lists channel handling: audio is mono for v1 and at most stereo for v2, and 5.1 is not preserved.
That is a packaged product: you upload a file and it separates speakers, translates and re-voices.
What Sume gives you
Sume does not sell a one-click dubbing product in the API reference. It sells the parts: POST /v1/stt-1.0/transcribe for timed text, POST /v1/tts-1.0/generate for speech, and Timeline audio to place and join clips. A TTS job renders one voice, selected by avatar_id, avatar_handle or voice.id.
STT does not expose a speaker field in the API reference: diarization options are fixed server-side. So if you need to know who is speaking, you must carry that from your own source, such as a script with speaker labels or separate tracks.
| Question | ElevenLabs dubbing | Sume pieces |
|---|---|---|
| Speakers per file | Up to 32 | Not a limit you set; one voice per TTS job |
| Detects who speaks | Part of the product | No speaker field in the STT docs |
| Assigns a voice per speaker | Product handles it | You split text per speaker and pick a voice for each |
| Channel handling | Mono v1, stereo max v2 | You choose channels mono or source in audio detach |
A multi-speaker plan on Sume
Work from labelled text. If you have a script, split it by speaker; each speaker's lines become jobs with that speaker's voice. Use sentence segments from STT to time the lines, then place each rendered clip with Timeline audio.
- Transcribe with
segmentation.mode: sentencefor gapless sentence times. - Group lines by speaker in your own code.
- Render each speaker's lines with a fixed voice.
- Concatenate or place clips with wav, and mux to the video at the end.
When to pick which
If you need automatic speaker handling on a long interview with many voices and have no script, the packaged tool does more for you. If you control the script, want exact per-line voice choices, or are building a repeatable pipeline you can retry safely with Idempotency-Key, composing pieces is predictable and each step is priced separately.
For the single-speaker case the difference mostly disappears.
Cost shape of a composed pipeline
A composed dub has four priced steps: audio detach at $0.01, transcription at $0.01 per audio minute, speech at $0.0475 per 1,000 characters, and Timeline audio at $0.01. For a 5-minute two-speaker clip with about 4,000 characters of translated text, that is roughly $0.01 + $0.05 + $0.19 + $0.01 = $0.26, before any translation step you run yourself.
Translation is not a Sume endpoint in the pieces above, so bring your own and count its cost separately.
Takeaway
ElevenLabs advertises 32 speakers per file as a feature. On Sume the equivalent is your own speaker split, one voice per TTS job. See the dubbing pipeline walkthrough for the steps.
Sources
Related posts
More in Comparisons
- ElevenLabs Dubbing v2 skips the transcript: Sume steps you can read
ElevenLabs says Dubbing v2 conditions on the source performance, not a transcript. Sume dubs run through text you can read and fix before any voice is made.
- ElevenLabs Opus 48 kHz output vs Sume TTS mp3, wav and raw formats
ElevenLabs lists Opus, MP3, PCM and mu-law outputs. Sume TTS 1.0 offers mp3, wav and raw with six sample rates. See which fits your pipeline.
- ElevenLabs previous_text and next_text vs Sume TTS: split scripts
ElevenLabs lets you pass previous_text and next_text for prosody continuity. Sume TTS has no such fields, so here is how to keep long scripts smooth.
- ElevenLabs TTS seed and free regenerations vs Sume TTS retry cost
ElevenLabs offers a seed and up to two free regenerations of identical requests. On Sume a retry with the same Idempotency-Key returns the same job. Compare.
Written by Sume