Dub a 3-minute explainer into 4 languages: $0.43 per language

One 3-minute explainer dubbed into four languages costs up to $1.76 on Sume: $0.04 once to transcribe, then about $0.43 per language to speak and render.

5 min readSume
All posts

Dubbing a 3-minute explainer into four languages costs up to $1.76 on Sume: $0.04 once to detach and transcribe, then up to $0.43 per language for speech, a concat and a three-minute timeline render. The transcript is reused, and Index-Translate's listed free API does each translation.

Assumptions: 850 characters per minute of speech, so 2,550 characters per language, and a spine that stays at or under 180 seconds. The render at $0.10 per output minute is the biggest line. Rates are from the pricing page and the Timeline docs, read 2026-10-05.

Once, then per language

Transcribe once, then repeat only the language-specific steps.

Up to cost of a 3-minute dub by number of languages at list rates (read 2026-10-05)
LanguagesShared (detach + STT)Per languageTotal, up to
1$0.04$0.43$0.47
2$0.04$0.43$0.90
4$0.04$0.43$1.76
8$0.04$0.43$3.49

Where to save

  • Transcribe once. Speech-to-text takes up to 10 minutes of audio per request, so a 3-minute video is one request.
  • Use the unbilled POST /v1/timeline-1.0/plan to see billable_minutes before you render.
  • A language whose translation speaks longer than 180 seconds adds a billable minute, so tighten that script.

Render one language

This is the last step for one language: concat the speech for its length and render it under the video.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def run(path, body, key):
    r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
    r.raise_for_status()
    job = r.json()["data"]
    while not requests.get(job["status_url"], headers=H, timeout=30).json()["data"]["terminal"]:
        time.sleep(2)
    res = requests.get(job["result_url"], headers=H, timeout=30)
    res.raise_for_status()
    return res.json()["data"]["result"]

VIDEO, SPEECH = os.environ["VIDEO_URL"], os.environ["SPEECH_URL"]   # both on media.sume.com
joined = run("/v1/timeline-1.0/audio", {"operation": "concat", "parts": [{"url": SPEECH}]}, "dub-0412-de-spine")
secs = joined["duration_seconds"]
body = {"audio": {"url": joined["audio_url"], "duration_seconds": secs},
        "video": [{"source_url": VIDEO, "start": 0, "duration": secs}]}
out = run("/v1/timeline-1.0/render", body, "dub-0412-de-render")
print(out["video_url"])

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume