Dubbing a 3-minute video on Sume: 46 cents before captions

Sume's docs list no dubbing endpoint. You assemble one: STT 3 cents, your translation, TTS 13 cents and a 30-cent render come to 46 cents before captions.

5 min readSume
All posts

What exists and what does not

Sume's docs do not describe a dedicated dubbing endpoint, so a dub is a pipeline you assemble from parts that are documented: transcribe the speech, translate the text with your own tool, synthesize the new language with TTS, and put the new audio under the picture with a Timeline render. Translation is outside Sume's billing in this plan.

One caution from the docs: video models do not lip-sync to generated TTS or a later voice-over, so a close-up talking face will not match the new audio.

Cost of the Sume steps for 3 minutes

Assume 2,700 characters of translated script (an assumption: 900 characters per minute). 2,700 x 0.00475 = 12.83 cents, rounded up to 13.

Dubbing a 3-minute video, Sume line items (rates read 2026-10-08)
StepBasisCost
Detach audio and transcribe (STT)3 audio minutes x $0.013 cents (plus free probe)
TranslateYour own toolNot billed by Sume
TTS in the new language2,700 chars x 0.00475, rounded up13 cents
CaptionsOver 60 s: check live price; $0.20 is for up to 60 sNot included
Timeline renderceil(3) x $0.1030 cents
Sume subtotal without captions46 cents

Captions are the open line

The fixed $0.20 estimate is for videos up to 60 seconds. A 3-minute video needs the live price from GET /v1/catalog. If you instead caption three 60-second segments, 3 x $0.20 = 60 cents, which brings the total to 46 + 60 = 106 cents. The 46-cent subtotal excludes captions, so add the caption price you are quoted.

Order of work

Detach the audio once, transcribe, translate, then generate TTS in parts under the 20,000-character cap, join them if needed, and render with the joined audio as the spine. Pin a concrete TTS model id and copy the recorded voice settings between parts. References: Audio detach, Video inspect, Timeline 1.0, Jobs and results.

Timing the new audio

Translated speech is often longer or shorter than the original. Measure the new TTS duration before you render, because Timeline's audio.duration_seconds sets the output length and the video slots must fill it. If the dub is longer, you either trim the script or allow the video to hold or loop, which Timeline reports as soft warnings.

Keep cutting points aligned to scene changes, since word-for-word sync is not possible with a different language.

What this plan leaves out

Voice cloning of the original speaker, lip-sync and translation quality are not covered by the documented Sume surfaces used here. For faceless explainers, product demos and screen recordings the approach works well; for on-camera speakers, plan on subtitles rather than a dub.

Sources

More in Use cases

All Use cases posts

Written by Sume