Dubbing a 3-minute video on Sume: 46 cents before captions
Sume's docs list no dubbing endpoint. You assemble one: STT 3 cents, your translation, TTS 13 cents and a 30-cent render come to 46 cents before captions.

What exists and what does not
Sume's docs do not describe a dedicated dubbing endpoint, so a dub is a pipeline you assemble from parts that are documented: transcribe the speech, translate the text with your own tool, synthesize the new language with TTS, and put the new audio under the picture with a Timeline render. Translation is outside Sume's billing in this plan.
One caution from the docs: video models do not lip-sync to generated TTS or a later voice-over, so a close-up talking face will not match the new audio.
Cost of the Sume steps for 3 minutes
Assume 2,700 characters of translated script (an assumption: 900 characters per minute). 2,700 x 0.00475 = 12.83 cents, rounded up to 13.
| Step | Basis | Cost |
|---|---|---|
| Detach audio and transcribe (STT) | 3 audio minutes x $0.01 | 3 cents (plus free probe) |
| Translate | Your own tool | Not billed by Sume |
| TTS in the new language | 2,700 chars x 0.00475, rounded up | 13 cents |
| Captions | Over 60 s: check live price; $0.20 is for up to 60 s | Not included |
| Timeline render | ceil(3) x $0.10 | 30 cents |
| Sume subtotal without captions | 46 cents |
Captions are the open line
The fixed $0.20 estimate is for videos up to 60 seconds. A 3-minute video needs the live price from GET /v1/catalog. If you instead caption three 60-second segments, 3 x $0.20 = 60 cents, which brings the total to 46 + 60 = 106 cents. The 46-cent subtotal excludes captions, so add the caption price you are quoted.
Order of work
Detach the audio once, transcribe, translate, then generate TTS in parts under the 20,000-character cap, join them if needed, and render with the joined audio as the spine. Pin a concrete TTS model id and copy the recorded voice settings between parts. References: Audio detach, Video inspect, Timeline 1.0, Jobs and results.
Timing the new audio
Translated speech is often longer or shorter than the original. Measure the new TTS duration before you render, because Timeline's audio.duration_seconds sets the output length and the video slots must fill it. If the dub is longer, you either trim the script or allow the video to hold or loop, which Timeline reports as soft warnings.
Keep cutting points aligned to scene changes, since word-for-word sync is not possible with a different language.
What this plan leaves out
Voice cloning of the original speaker, lip-sync and translation quality are not covered by the documented Sume surfaces used here. For faceless explainers, product demos and screen recordings the approach works well; for on-camera speakers, plan on subtitles rather than a dub.
Sources
More in Use cases
- eBay listing clips from cutouts: remove background, then animate
For 20 eBay items, remove each photo's background at $0.0225, then make a 5-second clip with Wan 3.0. A request pair, a cost table, and a fidelity check.
- Esports team intro in 21:9: Seedance 2.5 at 720p, 8 s, with cost
Seedance 2.5 on Sume renders 21:9 at 1470 x 630 for 720p. An 8-second stream-banner intro costs $4.65 at 720p, $2.15 at 480p. Rates and limits.
- Esports team intro sting: 5 seconds with a logo reference
A 5-second team intro for a stream in 16:9: send the logo as a reference image to Omni Flash, $0.625 at 720p on Sume, and fix the text in a short edit.
- Etsy-style listing gallery: 10 images per product, cost by image model
Ten gallery images for one product cost $0.07 to $1.00 on Sume depending on the model; 25 products is 250 images. Table with seedream, flux and gpt-image-2.5.
Written by Sume