Dub a 3-minute explainer into 4 languages: $0.43 per language
One 3-minute explainer dubbed into four languages costs up to $1.76 on Sume: $0.04 once to transcribe, then about $0.43 per language to speak and render.

Dubbing a 3-minute explainer into four languages costs up to $1.76 on Sume: $0.04 once to detach and transcribe, then up to $0.43 per language for speech, a concat and a three-minute timeline render. The transcript is reused, and Index-Translate's listed free API does each translation.
Assumptions: 850 characters per minute of speech, so 2,550 characters per language, and a spine that stays at or under 180 seconds. The render at $0.10 per output minute is the biggest line. Rates are from the pricing page and the Timeline docs, read 2026-10-05.
Once, then per language
Transcribe once, then repeat only the language-specific steps.
| Languages | Shared (detach + STT) | Per language | Total, up to |
|---|---|---|---|
| 1 | $0.04 | $0.43 | $0.47 |
| 2 | $0.04 | $0.43 | $0.90 |
| 4 | $0.04 | $0.43 | $1.76 |
| 8 | $0.04 | $0.43 | $3.49 |
Where to save
- Transcribe once. Speech-to-text takes up to 10 minutes of audio per request, so a 3-minute video is one request.
- Use the unbilled
POST /v1/timeline-1.0/planto seebillable_minutesbefore you render. - A language whose translation speaks longer than 180 seconds adds a billable minute, so tighten that script.
Render one language
This is the last step for one language: concat the speech for its length and render it under the video.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
r.raise_for_status()
job = r.json()["data"]
while not requests.get(job["status_url"], headers=H, timeout=30).json()["data"]["terminal"]:
time.sleep(2)
res = requests.get(job["result_url"], headers=H, timeout=30)
res.raise_for_status()
return res.json()["data"]["result"]
VIDEO, SPEECH = os.environ["VIDEO_URL"], os.environ["SPEECH_URL"] # both on media.sume.com
joined = run("/v1/timeline-1.0/audio", {"operation": "concat", "parts": [{"url": SPEECH}]}, "dub-0412-de-spine")
secs = joined["duration_seconds"]
body = {"audio": {"url": joined["audio_url"], "duration_seconds": secs},
"video": [{"source_url": VIDEO, "start": 0, "duration": secs}]}
out = run("/v1/timeline-1.0/render", body, "dub-0412-de-render")
print(out["video_url"])
Sources
Related posts
More in Use cases
- Dub a 45-second ad into Spanish and German: which steps Sume covers
Detach, transcribe, speak, join and caption are Sume jobs worth about 70 cents for two languages. Translation is not on the pages we read; bring your own.
- Dub a 60-second promo into German: step prices, up to $0.17
A 60-second promo dubbed into German costs up to $0.17 on Sume: detach, speech-to-text, TTS for 900 characters, a concat and a one-minute timeline render.
- Dubbing a video with music: what replacing the audio track means
A dub built from STT, TTS and a timeline swaps the whole audio track. The original music goes with it unless you add a bed back. Options on Sume, step by step.
- A dubbed line runs 40% longer: TTS speed tops out at 1.5, so rewrite
Translated lines often run long. Sume TTS speed is 0.6 to 1.5, so a line 40% too long can be sped up, but one 60% too long must be rewritten.
Written by Sume