AI dubbing rate per minute: a step-by-step cost breakdown

An AI dub's rate per minute is the sum of its steps: transcribe, translate, speak, render. Sume's per-step prices and a worked 1- and 10-minute dub.

5 min readSume
All posts

An AI dubbing rate per minute is the sum of the steps a dub takes: transcribe the original speech, translate it, speak the translation, and put the new voice under the video. On Sume those steps are metered separately, and with a 900-character translated script per minute a 10-minute dub works out to at most $0.15275 per minute before the 5.5% agent fee, translation not included.

Rates are read from the code behind API pricing. Limits come from the TTS schema in the Sume API reference (the OpenAPI document behind the API reference docs), Video inspect and Timeline 1.0, read on 2026-09-29. Translate a video's voiceover by API has the requests for each step.

What does one minute of AI dubbing cost?

The script length is an assumption: 900 characters of translated text per minute of video. Count your own translated script, because speech is billed per character, spaces and punctuation included. A 1-minute dub comes to at most $0.15275. Probing the clip is free; only its transcript is billed.

Computed from API pricing, Video inspect and Timeline 1.0, read 2026-09-29. Before the 5.5% agent fee.
StepRate1-minute video10-minute video
Transcribe (video inspect, transcribe: true)$0.01 per audio minute$0.01$0.1
TranslateYour own stepNot a Sume chargeNot a Sume charge
Speak the translation$0.0475 per 1,000 characters$0.04275$0.4275
Render video over the new voiceUp to $0.10 per output minuteUp to $0.1Up to $1
TotalUp to $0.15275Up to $1.5275

Why is the render an "up to" price?

Timeline 1.0 reserves its rate per started output minute. In the current pricing code the render then captures its own compute cost, never above that reservation, so the row shows the ceiling. Failed jobs release or refund their reservation where applicable.

What limits change the math for long videos?

  • The video must already be a file in your workspace on media.sume.com, such as an earlier Sume job's output: video inspect and Timeline read nothing else.
  • The transcript reserves by duration_seconds, at most 600 seconds; omit it and 1 minute is reserved. A clip with no audio track fails with inspect_source_has_no_audio.
  • One TTS request takes up to 20,000 characters and fails with tts_duration_exceeded, with no credit captured, past 1,200 seconds of audio.
  • Set TTS language for every non-English transcript; omitted, it defaults to English.
  • A Timeline render outputs 1 to 1,800 seconds.

What does this price not include?

  • Translation. You translate the lines yourself or with another tool, and that cost sits outside Sume's rates.
  • The original music and effects. In the current code a render plays only its audio spine and an optional soundtrack, so each clip's own audio is dropped. Supply a Sume-hosted bed as soundtrack if you have one.
  • Lip sync. Sume's model docs say video models do not lip-sync to a later voice-over, so a speaker on camera won't match the new voice; Lip sync vs dubbing explains the difference.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume