ElevenLabs dubbing editor in maintenance mode: redo one line on Sume

ElevenLabs says its dubbing editor gets critical fixes only. If you need per-line regeneration, a Sume TTS job per sentence gives you one retry per line.

5 min readSume
All posts

ElevenLabs states that its Dubbing Studio, the editor with transcript editing, speaker reassignment and per-clip regeneration, is in maintenance mode and receives critical bug fixes only. If your workflow depends on redoing one line of a dub, the Sume answer is a pipeline rather than an editor: one TTS job per sentence, so a bad line is one retry, joined at the end with a gapless concat.

The ElevenLabs facts are from its Dubbing documentation (read 2026-10-10). Sume's side is from Audio detach, Timeline audio and the Sume OpenAPI contract.

What ElevenLabs says about its dubbing paths

The page describes automatic dubbing and a studio editor. Automatic dubbing accepts up to 1 GB and 180 minutes in the app, or 3 GB per source file via the API; the editor accepts up to 1 GB and 45 minutes. It lists 90+ languages, up to 32 unique speakers per file, and voice cloning with a cloning-strength setting. Creating and downloading dubs is available on all plans, while transcript editing and regeneration via the API are Enterprise-only.

Dubbing editing paths (ElevenLabs read 2026-10-10; Sume per docs.sume.com and OpenAPI)
NeedElevenLabsSume pipeline
Redo one linePer-clip regeneration in the editor (maintenance mode); via API on EnterpriseRe-run one TTS job for that sentence
Edit the translationTranscript editing in the editorEdit your own script text, then submit it
Join the linesHandled inside the dubTimeline audio concat, up to 20 parts, $0.01 per job
SpeakersUp to 32 unique speakers per fileOne voice per TTS job; a voice per speaker

The Sume line-by-line pipeline

Detach the audio from the source video once ($0.01 per job, source up to 1,800 seconds, output up to 900 seconds per detach). Transcribe it with Sume STT to get word timings and, with segmentation.mode: "sentence", gapless sentence segments. Translate the sentences yourself, then submit one TTS job per sentence with the target language set. A TTS job is priced at $0.0475 per 1,000 characters, rounded up to a whole cent with a one-cent minimum per job (the catalog formula is list times 1.25, ceiling to cents), so any sentence under about 210 characters costs one cent.

When one line sounds wrong, you resubmit that one sentence with a new idempotency key and replace its part in the concat. The other lines are untouched and already paid for. That is the closest documented equivalent to per-clip regeneration.

Join with WAV and keep the order

Ask TTS for WAV output when you plan to join, because the Timeline audio page says MP3 adds priming padding at every edge. Concat accepts 1 to 20 parts per job, so a 60-line dub needs three concat jobs, then one more to join the three results. All parts must share a channel layout, or the job fails with audio_parts_channel_mismatch.

Total cost for a 60-sentence dub with sentences of 100 characters each: 60 TTS jobs at the one-cent minimum is $0.60, plus four $0.01 concat jobs, which is $0.64 before STT and detach. Redoing three bad lines adds three more one-cent TTS jobs and the re-concat, about $0.05 in all if you re-join one affected group of 20 and then the final join.

What this does not replace

It does not replace speaker diarization, lip-sync to the original timing or voice cloning of the original speakers. Sume's STT fixes its diarize flag server-side, and voice cloning is app-only, not an API feature. If you need an automatic multi-speaker dub with cloning, use the vendor that lists it. If your goal is a controllable, line-addressable dub, the pipeline above is the route Sume documents.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume