ElevenLabs dubbing 32 speakers per file vs Sume one voice per TTS job

ElevenLabs dubbing handles up to 32 speakers in a file. Sume composes dubbing from STT and TTS, one voice per job. What that means for multi-speaker video.

6 min readSume
All posts

What the vendor page says

The ElevenLabs dubbing page says a file can contain up to 32 speakers and that dubbing covers 90+ languages. It also lists channel handling: audio is mono for v1 and at most stereo for v2, and 5.1 is not preserved.

That is a packaged product: you upload a file and it separates speakers, translates and re-voices.

What Sume gives you

Sume does not sell a one-click dubbing product in the API reference. It sells the parts: POST /v1/stt-1.0/transcribe for timed text, POST /v1/tts-1.0/generate for speech, and Timeline audio to place and join clips. A TTS job renders one voice, selected by avatar_id, avatar_handle or voice.id.

STT does not expose a speaker field in the API reference: diarization options are fixed server-side. So if you need to know who is speaking, you must carry that from your own source, such as a script with speaker labels or separate tracks.

Multi-speaker dubbing, packaged vs composed (read 2026-10-03)
QuestionElevenLabs dubbingSume pieces
Speakers per fileUp to 32Not a limit you set; one voice per TTS job
Detects who speaksPart of the productNo speaker field in the STT docs
Assigns a voice per speakerProduct handles itYou split text per speaker and pick a voice for each
Channel handlingMono v1, stereo max v2You choose channels mono or source in audio detach

A multi-speaker plan on Sume

Work from labelled text. If you have a script, split it by speaker; each speaker's lines become jobs with that speaker's voice. Use sentence segments from STT to time the lines, then place each rendered clip with Timeline audio.

  • Transcribe with segmentation.mode: sentence for gapless sentence times.
  • Group lines by speaker in your own code.
  • Render each speaker's lines with a fixed voice.
  • Concatenate or place clips with wav, and mux to the video at the end.

When to pick which

If you need automatic speaker handling on a long interview with many voices and have no script, the packaged tool does more for you. If you control the script, want exact per-line voice choices, or are building a repeatable pipeline you can retry safely with Idempotency-Key, composing pieces is predictable and each step is priced separately.

For the single-speaker case the difference mostly disappears.

Cost shape of a composed pipeline

A composed dub has four priced steps: audio detach at $0.01, transcription at $0.01 per audio minute, speech at $0.0475 per 1,000 characters, and Timeline audio at $0.01. For a 5-minute two-speaker clip with about 4,000 characters of translated text, that is roughly $0.01 + $0.05 + $0.19 + $0.01 = $0.26, before any translation step you run yourself.

Translation is not a Sume endpoint in the pieces above, so bring your own and count its cost separately.

Takeaway

ElevenLabs advertises 32 speakers per file as a feature. On Sume the equivalent is your own speaker split, one voice per TTS job. See the dubbing pipeline walkthrough for the steps.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume