ElevenLabs Dubbing v2 API: what it does, and what Sume offers
ElevenLabs Dubbing v2 is a project-based API with editable transcripts. Sume has no dubbing route; here is how its STT, TTS and audio-join steps map on.

ElevenLabs' changelog lists a Dubbing v2 API, added on August 10, 2026, that works on a project, covers 90+ languages, exposes editable JSON transcripts and translations, and regenerates only the segments you change. Sume does not list a dubbing route. It lists the parts a dub is built from, speech-to-text, text-to-speech and audio joins, and you hold the project state yourself.
What did ElevenLabs say Dubbing v2 does?
These points come from ElevenLabs' own changelog entry, not from tests we ran. It describes the API as project-based, with 90+ languages, editable JSON transcripts and translations, and selective regeneration of changed segments. The entry does not give prices or limits, so none are stated here.
Which Sume steps map onto each part?
The mapping is a rough one: Sume's steps are separate jobs, and the project is your own record of them.
| Dubbing v2 feature | Closest Sume route | Who holds the state |
|---|---|---|
| Editable transcript | POST /v1/stt-1.0/transcribe with sentence segments | You store the segments and edit them |
| Translations | None on Sume's dubbing side | You translate in your own code |
| Regenerate a changed segment | POST /v1/tts-1.0/generate for that sentence | You keep one audio file per sentence |
| Assembled dub track | POST /v1/timeline-1.0/audio with operation: "concat" | The job returns one file and offsets |
How do I redo one changed sentence?
Speak the edited sentence again and run the join again with the new file in its place. Timeline audio joins in the sample domain, with no re-synthesis and no silence at the seams, so the untouched files are not spoken twice. A join takes 1 to 20 parts, so a track with more sentences needs its joins planned in stages. Detail on offsets is in the Timeline audio docs.
What is the project on Sume?
A Sume dub has no project object, so make your own. A small record is enough: the source video URL, the list of sentences with their start and end times, the translated text for each, the job id and audio URL of each spoken sentence, and the URL of the last joined track. Editing a translation then means changing one row, re-running one text-to-speech job and one join. Each job takes an idempotency key where the docs require one, so a retry does not create a second paid job.
If you need the language coverage or editor that ElevenLabs describes, use ElevenLabs' own product for that. Sume's part is the media handling around it: Sume-hosted audio in, Sume-hosted audio out, and durable media.sume.com URLs.
What does the Sume side cost?
Extraction is $0.01 per detach job and the join is $0.01 flat per job, with no provider inference on either. Speech-to-text is $0.01 per audio minute. Text-to-speech is billed by character, so a small edit costs only the characters you re-speak. The speaker's lips are unchanged by any of these steps; see lip sync versus dubbing for that gap.
Sources
Related posts
More in Developers
- ElevenLabs music API: composition plans vs a prompt-only API
ElevenLabs added Music v2.5 API support on 2026-09-14 with composition plans. How a plan differs from Sume's prompt-only music route with no duration field
- Enterprise API rate limit: what a Sume key gets by default
Sume's Enterprise rate limit is contracted, not self-serve. Until a number is provisioned, an Enterprise key gets the Scale row, plus 20 concurrent jobs.
- Export API usage to CSV: turning Sume's usage ledger into rows
Sume's docs list no CSV export, but GET /v1/usage returns ledger rows as JSON. Convert them with jq, and keep captured rows separate from refunds.
- Retry a failed run: same idempotency key returns old failure
Retry a failed run: the same Sume Idempotency-Key returns the old failure. Use a new key, or continue with previous_run_id to keep finished clips.
Written by Sume