ElevenLabs Dubbing v2 skips the transcript: Sume steps you can read

ElevenLabs says Dubbing v2 conditions on the source performance, not a transcript. Sume dubs run through text you can read and fix before any voice is made.

5 min readSume
All posts

ElevenLabs describes Dubbing v2 as an audio-to-audio model that conditions on the source performance, not on a transcript. That removes the text stage where you would normally catch a mistranslated product name. A Sume dub keeps that stage: you read the transcript and the translation before any voice is generated.

The ElevenLabs details come from its Dubbing Studio page, read on 2026-10-03. It lists 90+ languages and accents, automatic voice cloning that preserves speaker identity, and plans from Free to Pro.

What does the ElevenLabs page claim?

The page says Dubbing v2 covers 90+ languages and accents, clones the original speaker's voice automatically across target languages, and is fully automated with no manual pipeline. It names two access routes: ElevenCreative for one-click dubbing and ElevenProductions for broadcast-quality work with human translators and professional mixing.

Monthly-billing plans show how many Dubbing v2 minutes each includes. Extra minutes run $2.23 to $4.91 depending on plan tier.

ElevenLabs Dubbing v2 plan minutes, monthly billing (read 2026-10-03)
PlanMonthly priceCreditsDubbing v2 minutes
Free$010,0000.4
Starter$6 ($1 first month)30,0002
Creator$22 ($11 first month)121,0009
Pro$99600,00044

Why does skipping the transcript matter?

A transcript is a checkpoint. Brand names, numbers and legal lines are the content most likely to be translated wrongly, and they are also the content a reviewer would check first. In an audio-to-audio flow, the first time anyone hears the error is in the finished dub.

That is a trade, not a fault. Audio-to-audio systems can carry timing, emphasis and emotion that a text hop loses. The page's claim is about performance; the cost is the missing review point. Decide which risk you can afford for the asset in hand.

What is the reviewable path on Sume?

Every step is a job with its own result. Detach the audio with audio detach, transcribe with video inspect (transcribe: true, segmentation.mode: "sentence"), translate the sentences with a translator you choose, voice each line, and join with timeline audio. The first two and last are $0.01 per job or audio minute per their pages; translation and speech are priced separately.

You can stop after the transcript. Read the sentences, fix the brand name, and only then spend on voice. The transcript is a text field in the job result, so a script or a person can diff it.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: review-stt-001" \
  -d '{"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "frames": false, "transcribe": true,
       "segmentation": {"mode": "sentence"}}'

Which should you choose for which asset?

Choose by the cost of an error, not by feature lists. A 30-second social teaser with no numbers can ride on an automated dub. A pricing explainer or a regulated claim should pass through text a person reads.

  • Speaker identity matters and review does not: an audio-to-audio dubbing tool fits.
  • Wording must be exact: keep a transcript checkpoint, as in the Sume chain.
  • You need burned subtitles as well: Sume's video captions are $0.20 per job up to 60 seconds.
  • You need the speaker's own voice in 90 languages: Sume's docs do not offer that.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume