ElevenLabs Dubbing v2 API: what it does, and what Sume offers

ElevenLabs Dubbing v2 is a project-based API with editable transcripts. Sume has no dubbing route; here is how its STT, TTS and audio-join steps map on.

4 min readSume
All posts

ElevenLabs' changelog lists a Dubbing v2 API, added on August 10, 2026, that works on a project, covers 90+ languages, exposes editable JSON transcripts and translations, and regenerates only the segments you change. Sume does not list a dubbing route. It lists the parts a dub is built from, speech-to-text, text-to-speech and audio joins, and you hold the project state yourself.

What did ElevenLabs say Dubbing v2 does?

These points come from ElevenLabs' own changelog entry, not from tests we ran. It describes the API as project-based, with 90+ languages, editable JSON transcripts and translations, and selective regeneration of changed segments. The entry does not give prices or limits, so none are stated here.

Which Sume steps map onto each part?

The mapping is a rough one: Sume's steps are separate jobs, and the project is your own record of them.

ElevenLabs Dubbing v2 features as described in its changelog, against Sume routes, read 2026-09-29.
Dubbing v2 featureClosest Sume routeWho holds the state
Editable transcriptPOST /v1/stt-1.0/transcribe with sentence segmentsYou store the segments and edit them
TranslationsNone on Sume's dubbing sideYou translate in your own code
Regenerate a changed segmentPOST /v1/tts-1.0/generate for that sentenceYou keep one audio file per sentence
Assembled dub trackPOST /v1/timeline-1.0/audio with operation: "concat"The job returns one file and offsets

How do I redo one changed sentence?

Speak the edited sentence again and run the join again with the new file in its place. Timeline audio joins in the sample domain, with no re-synthesis and no silence at the seams, so the untouched files are not spoken twice. A join takes 1 to 20 parts, so a track with more sentences needs its joins planned in stages. Detail on offsets is in the Timeline audio docs.

What is the project on Sume?

A Sume dub has no project object, so make your own. A small record is enough: the source video URL, the list of sentences with their start and end times, the translated text for each, the job id and audio URL of each spoken sentence, and the URL of the last joined track. Editing a translation then means changing one row, re-running one text-to-speech job and one join. Each job takes an idempotency key where the docs require one, so a retry does not create a second paid job.

If you need the language coverage or editor that ElevenLabs describes, use ElevenLabs' own product for that. Sume's part is the media handling around it: Sume-hosted audio in, Sume-hosted audio out, and durable media.sume.com URLs.

What does the Sume side cost?

Extraction is $0.01 per detach job and the join is $0.01 flat per job, with no provider inference on either. Speech-to-text is $0.01 per audio minute. Text-to-speech is billed by character, so a small edit costs only the characters you re-speak. The speaker's lips are unchanged by any of these steps; see lip sync versus dubbing for that gap.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume