ElevenLabs Dubbing v2 skips the transcript: Sume steps you can read
ElevenLabs says Dubbing v2 conditions on the source performance, not a transcript. Sume dubs run through text you can read and fix before any voice is made.

ElevenLabs describes Dubbing v2 as an audio-to-audio model that conditions on the source performance, not on a transcript. That removes the text stage where you would normally catch a mistranslated product name. A Sume dub keeps that stage: you read the transcript and the translation before any voice is generated.
The ElevenLabs details come from its Dubbing Studio page, read on 2026-10-03. It lists 90+ languages and accents, automatic voice cloning that preserves speaker identity, and plans from Free to Pro.
What does the ElevenLabs page claim?
The page says Dubbing v2 covers 90+ languages and accents, clones the original speaker's voice automatically across target languages, and is fully automated with no manual pipeline. It names two access routes: ElevenCreative for one-click dubbing and ElevenProductions for broadcast-quality work with human translators and professional mixing.
Monthly-billing plans show how many Dubbing v2 minutes each includes. Extra minutes run $2.23 to $4.91 depending on plan tier.
| Plan | Monthly price | Credits | Dubbing v2 minutes |
|---|---|---|---|
| Free | $0 | 10,000 | 0.4 |
| Starter | $6 ($1 first month) | 30,000 | 2 |
| Creator | $22 ($11 first month) | 121,000 | 9 |
| Pro | $99 | 600,000 | 44 |
Why does skipping the transcript matter?
A transcript is a checkpoint. Brand names, numbers and legal lines are the content most likely to be translated wrongly, and they are also the content a reviewer would check first. In an audio-to-audio flow, the first time anyone hears the error is in the finished dub.
That is a trade, not a fault. Audio-to-audio systems can carry timing, emphasis and emotion that a text hop loses. The page's claim is about performance; the cost is the missing review point. Decide which risk you can afford for the asset in hand.
What is the reviewable path on Sume?
Every step is a job with its own result. Detach the audio with audio detach, transcribe with video inspect (transcribe: true, segmentation.mode: "sentence"), translate the sentences with a translator you choose, voice each line, and join with timeline audio. The first two and last are $0.01 per job or audio minute per their pages; translation and speech are priced separately.
You can stop after the transcript. Read the sentences, fix the brand name, and only then spend on voice. The transcript is a text field in the job result, so a script or a person can diff it.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: review-stt-001" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "frames": false, "transcribe": true,
"segmentation": {"mode": "sentence"}}'Which should you choose for which asset?
Choose by the cost of an error, not by feature lists. A 30-second social teaser with no numbers can ride on an automated dub. A pricing explainer or a regulated claim should pass through text a person reads.
- Speaker identity matters and review does not: an audio-to-audio dubbing tool fits.
- Wording must be exact: keep a transcript checkpoint, as in the Sume chain.
- You need burned subtitles as well: Sume's video captions are $0.20 per job up to 60 seconds.
- You need the speaker's own voice in 90 languages: Sume's docs do not offer that.
Sources
Related posts
More in Comparisons
- ElevenLabs v4 Turbo vs Flash vs v3: price per 1,000 characters
ElevenLabs API list per 1K characters: v4 Turbo $0.011 until Oct 12 (then $0.04), Flash/Turbo $0.04, v3 Multilingual $0.08, v4 $0.022 (then $0.08).
- Face and body swap video AI: Recast vs Sume's avatar face swap
Sume has two ways to put someone else in a video: H3 Max Recast swaps the person from a photo, Beta Face Swap applies a ready avatar's face. Which to use.
- Face swap vs H3 Max Recast: which swaps the person in a video?
Sume offers two ways to put a different person in a video: Avatar Face Swap (Beta) and H3 Max Recast. Inputs, length limits, audio and price side by side.
- fal list vs Sume price for Seedance 2.5: what 100 clips cost extra
For 100 five-second 720p Seedance 2.5 clips, fal list is about $231 and Sume is about $289: $57.78 more, the 25 percent house margin. What it covers.
Written by Sume