ElevenLabs Dubbing v2 cloning_strength 0-10: what Sume has instead
ElevenLabs Dubbing v2 has cloning_strength from 0 to 10, default 7. Sume has no such dial: you pick a voice per language and set speed. What that costs you.

ElevenLabs Dubbing v2 has a cloning_strength setting from 0 to 10 with a default of 7. Higher values favour resemblance to the original speaker; lower values give the model more freedom for natural delivery in the target language (ElevenLabs dubbing docs, read 2026-10-02). Sume has no equivalent dial: you choose a library voice and the language, and text to speech reads your translated script.
What the ElevenLabs page says
The docs label v2 as Alpha with automatic processing and no in-app transcript editor, and v1 as in maintenance mode with Dubbing Studio. Its dubbing API blog (read 2026-10-02) says Dubbing v2 does not adjust lips: audio lands where the original dialogue did through sync-aware translation.
| Item | What the docs say |
|---|---|
| cloning_strength | 0 to 10, default 7 |
| Higher value | Prioritises similarity to the original speaker |
| Lower value | More natural delivery in the target language |
| v2 languages | 90+ with regional dialects for some, such as en-AU |
| File limits | 1 GB and 180 minutes in the app; 3 GB per file via the API for v2 |
| Speakers | Up to 32 unique speakers per file |
| Self-serve concurrency | Up to 3 concurrent dubbing jobs |
| Background audio | Kept, so music and effects are not re-mixed |
What you control on Sume
On Sume you set the voice, the language and the text. Text to speech also takes speed, and sentence segments let you fit each line to its slot, as covered in Translate a video to another language with AI voice, on time. The voice does not shift toward the original speaker by a dial; it is simply the voice you picked.
Background audio is yours to keep: detach the original track, and mix the new speech over it with Timeline audio parts.
The trade-off
Which is better depends on whether the original speaker's identity matters more than control.
- ElevenLabs trades some naturalness for resemblance at high cloning_strength; Sume has no resemblance setting at all.
- ElevenLabs v2 handles speaker separation; on Sume you separate speakers yourself.
- Sume gives step-level control and per-step pricing; ElevenLabs gives one project per dub.
Sources
Related posts
More in Comparisons
- ElevenLabs maximum_text_length_per_request vs Sume max_characters
ElevenLabs now says to read maximum_text_length_per_request, not the old per-plan fields. Sume's TTS Router catalog publishes max_characters on every model row.
- ElevenLabs Music inpainting: redo one section vs a new Sume take
ElevenLabs Music v2 and v2.5 can regenerate one section of a song in the UI. Sume cannot edit a track: write a new brief for that part and generate again.
- ElevenLabs Music length: 3 seconds to 5 or 10 minutes vs Sume
ElevenLabs' two pages disagree on max Music length: 5 minutes on the capabilities page, 600000 ms on the Compose reference. Sume has no length field at all.
- ElevenLabs TTS output_format list vs Sume container and encoding
ElevenLabs names output_format strings like opus_48000_64 and ulaw_8000. Sume splits it into container, sample rate, bit rate and encoding, with no Opus.
Written by Sume