Runway voice dubbing API: 1 credit per 2 seconds vs Sume's dub steps

Runway prices voice dubbing at 1 credit per 2 seconds of audio: a 10-minute video is 300 credits. Sume has no dubbing endpoint, so it takes three steps.

5 min readSume
All posts

Runway lists its voice dubbing model, eleven_voice_dubbing, at 1 credit per 2 seconds of audio, and credits cost $0.01 each in the developer portal. That makes a 10-minute video 300 credits, or $3.00. Sume has no single dubbing call, so the same job is three steps: transcribe, translate the text yourself, then synthesize and join the audio.

Here are Runway's published numbers and request fields, and the Sume steps that replace the one call.

What does the Runway dubbing call cost and take?

The pricing page gives eleven_voice_dubbing at 1 credit per 2s of audio and says credits are $0.01 each. The unit is seconds of audio. The API reference lists the request fields as audioUri for the source file, targetLang for the language, numSpeakers, disableVoiceCloning, and dropBackgroundAudio.

The reference excerpt does not state a maximum length or the list of supported languages, so check the full reference before you plan a long file. Runway also sells voice isolation at 1 credit per 6 seconds and speech to speech at 1 credit per 3 seconds, which you can chain if you need cleaner stems first.

Runway dubbing cost by audio length at $0.01 per credit (read 2026-10-02)
Audio lengthCredits at 1 per 2 sCost
1 minute30$0.30
10 minutes300$3.00
30 minutes900$9.00
60 minutes1,800$18.00

Why does Sume not have a dubbing call?

Sume lists the building blocks, not a bundled dub. Speech to text, text to speech and the timeline are separate paid tools, and the hosted MCP exposes stt_create, tts_create and timeline_audio next to each other. Translation is the step in the middle, and it is text work your own agent or code does.

That makes the pipeline more flexible and less automatic. You choose the voice, you can edit the translated script line by line, and you re-render only the sentence you change. You also carry the speaker handling that Runway's numSpeakers field does for you: Sume's TTS takes one voice selector per request, so a multi-speaker dub is several requests.

What are the Sume steps in order?

  • Get the source transcript: video inspect with transcribe true runs Sume speech to text on a clip's audio and returns the transcript with the probe.
  • Translate the transcript in your own tool, keeping one sentence per line so timings stay mappable.
  • Synthesize each line with the TTS Router, using a ready avatar voice.
  • Join the audio parts with timeline audio concat, which takes up to 20 parts and is a flat per-job price with no re-synthesis.
  • Put the new audio on the original video through the Timeline, then check lip movement by eye, since dubbed audio does not re-time a talking face.

How do the costs compare?

Runway's cost is one multiplication on audio length. Sume's cost is the sum of three lines: transcription (billed per audio minute on video inspect, with one minute assumed when no duration hint is sent), synthesis per character at the provider list times 1.25, and a flat price for each concat job. Read the live numbers from the catalog before quoting a client.

For short clips, Runway's single call is simpler and the dollar difference is small. For long files, or where a client edits the translation, the stepwise route lets you change one sentence without paying for the whole file again. Whichever route you pick, voice cloning and dubbing of real people need consent and disclosure checks that sit outside the API.

Which should you pick?

Use Runway dubbing when you want a one-call dub with speaker handling and an optional switch to turn voice cloning off, and the clip is within its limits. Use Sume's steps when you control the translation, need a specific voice, or want to re-render single lines.

If you build the Sume route, test it on a 30-second clip first and keep each job id with its sentence so a re-render replaces exactly one segment. The video inspect page lists the transcribe fields and errors.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume