Dub a 60-second promo into German: step prices, up to $0.17

A 60-second promo dubbed into German costs up to $0.17 on Sume: detach, speech-to-text, TTS for 900 characters, a concat and a one-minute timeline render.

5 min readSume
All posts

Dubbing a 60-second promo into German costs up to $0.173 on Sume at list rates: $0.01 to detach the audio, $0.01 for one speech-to-text minute, $0.0428 to speak 900 characters, $0.01 to concat the speech so you can read its length, and $0.10 for a one-minute timeline render. Index-Translate does the translation on its listed free public API.

The 900 characters is an assumption for a minute of speech; use the length of your real translation. Prices are from the pricing page, Audio detach docs, Timeline audio docs and Timeline docs, read 2026-10-05. Render and concat are captured at their own compute, never above the reservation, so this is a ceiling.

The steps and what each costs

The render replaces the original audio with your new spine, so the original voice and any music on that track are gone.

Step prices for a 60-second dub at list rates (read 2026-10-05)
StepCountRateCost
Detach the audio of the video1$0.01 per job$0.01
Speech-to-text1 min$0.01 per minute$0.01
Translate (Index-Translate public API)1Listed as free$0.00
Sume TTS 1.0900 characters$0.0475 per 1,000$0.0428
Concat the speech for its length1$0.01 per job$0.01
Timeline render1 min$0.10 per output minute$0.10
Total, up to$0.173

Keep the spine under 60 seconds

Timeline reserves ceil(audio.duration_seconds / 60) minutes. A German translation that speaks for 61 seconds is a two-minute render at $0.20, so check the length from the concat result before you render and shorten the script if it is just over.

Render step

The earlier steps are the ones in the dubbing pipeline post. This runs the last two: it joins the speech artifact for a length and renders it under the original video, using the job flow from the jobs docs.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def run(path, body, key):
    r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
    r.raise_for_status()
    job = r.json()["data"]
    while not requests.get(job["status_url"], headers=H, timeout=30).json()["data"]["terminal"]:
        time.sleep(2)
    res = requests.get(job["result_url"], headers=H, timeout=30)
    res.raise_for_status()
    return res.json()["data"]["result"]

VIDEO, SPEECH = os.environ["VIDEO_URL"], os.environ["SPEECH_URL"]   # both on media.sume.com
joined = run("/v1/timeline-1.0/audio", {"operation": "concat", "parts": [{"url": SPEECH}]}, "dub-0412-de-spine")
secs = joined["duration_seconds"]
body = {"audio": {"url": joined["audio_url"], "duration_seconds": secs},
        "video": [{"source_url": VIDEO, "start": 0, "duration": secs}]}
out = run("/v1/timeline-1.0/render", body, "dub-0412-de-render")
print(out["video_url"])

Match the voice to the language

Speak German with a German voice. A voice whose primary language differs from language returns 409 tts_voice_language_mismatch before any job is created or charged.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume