Dub a 60-second promo into German: step prices, up to $0.17
A 60-second promo dubbed into German costs up to $0.17 on Sume: detach, speech-to-text, TTS for 900 characters, a concat and a one-minute timeline render.

Dubbing a 60-second promo into German costs up to $0.173 on Sume at list rates: $0.01 to detach the audio, $0.01 for one speech-to-text minute, $0.0428 to speak 900 characters, $0.01 to concat the speech so you can read its length, and $0.10 for a one-minute timeline render. Index-Translate does the translation on its listed free public API.
The 900 characters is an assumption for a minute of speech; use the length of your real translation. Prices are from the pricing page, Audio detach docs, Timeline audio docs and Timeline docs, read 2026-10-05. Render and concat are captured at their own compute, never above the reservation, so this is a ceiling.
The steps and what each costs
The render replaces the original audio with your new spine, so the original voice and any music on that track are gone.
| Step | Count | Rate | Cost |
|---|---|---|---|
| Detach the audio of the video | 1 | $0.01 per job | $0.01 |
| Speech-to-text | 1 min | $0.01 per minute | $0.01 |
| Translate (Index-Translate public API) | 1 | Listed as free | $0.00 |
| Sume TTS 1.0 | 900 characters | $0.0475 per 1,000 | $0.0428 |
| Concat the speech for its length | 1 | $0.01 per job | $0.01 |
| Timeline render | 1 min | $0.10 per output minute | $0.10 |
| Total, up to | $0.173 |
Keep the spine under 60 seconds
Timeline reserves ceil(audio.duration_seconds / 60) minutes. A German translation that speaks for 61 seconds is a two-minute render at $0.20, so check the length from the concat result before you render and shorten the script if it is just over.
Render step
The earlier steps are the ones in the dubbing pipeline post. This runs the last two: it joins the speech artifact for a length and renders it under the original video, using the job flow from the jobs docs.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
r.raise_for_status()
job = r.json()["data"]
while not requests.get(job["status_url"], headers=H, timeout=30).json()["data"]["terminal"]:
time.sleep(2)
res = requests.get(job["result_url"], headers=H, timeout=30)
res.raise_for_status()
return res.json()["data"]["result"]
VIDEO, SPEECH = os.environ["VIDEO_URL"], os.environ["SPEECH_URL"] # both on media.sume.com
joined = run("/v1/timeline-1.0/audio", {"operation": "concat", "parts": [{"url": SPEECH}]}, "dub-0412-de-spine")
secs = joined["duration_seconds"]
body = {"audio": {"url": joined["audio_url"], "duration_seconds": secs},
"video": [{"source_url": VIDEO, "start": 0, "duration": secs}]}
out = run("/v1/timeline-1.0/render", body, "dub-0412-de-render")
print(out["video_url"])
Match the voice to the language
Speak German with a German voice. A voice whose primary language differs from language returns 409 tts_voice_language_mismatch before any job is created or charged.
Sources
Related posts
More in Use cases
- Dubbing a video with music: what replacing the audio track means
A dub built from STT, TTS and a timeline swaps the whole audio track. The original music goes with it unless you add a bed back. Options on Sume, step by step.
- A dubbed line runs 40% longer: TTS speed tops out at 1.5, so rewrite
Translated lines often run long. Sume TTS speed is 0.6 to 1.5, so a line 40% too long can be sped up, but one 60% too long must be rewritten.
- E-bike Black Friday video ad: start and end frame reveal, three trims
An e-bike Black Friday ad from two stills: a 10-second start/end-frame clip on Gemini Omni Flash, then three trimmed cutdowns, about $1.51 in Sume jobs.
- Eleven v4 says 75% prefer it: run your own 20-listener test
ElevenLabs reports ~75% preference for Eleven v4 in blind tests. A pre-named voice winning 15 of 20 is ~2% by chance. Run your own test on Sume.
Written by Sume