Dub one Short into 8 languages: Python fan-out and the total cost
Detach and transcribe once, then run one TTS job and one render per language. A Python fan-out and the per-Short bill, from Sume's catalog rates.

To dub one short video into eight languages, do the shared work once and the per-language work eight times. Detach the audio and transcribe it one time, translate the script yourself, then submit one TTS job and one render per language. On Sume's published rates a 60-second Short comes to roughly $1.16 for eight languages with no captions, and about $2.76 with a burned-caption pass for each.
Rates come from the Sume catalog and docs pages for audio detach, video inspect, Timeline 1.0 and video captions, read 2026-10-03. The character count is an assumption you should replace with your own.
Which steps run once and which run per language?
The detach, transcript and render calls need the video on a media.sume.com URL, so import first. The reference puts the transcript on video inspect with transcribe: true, or on POST /v1/stt-1.0/transcribe with a detached audio URL.
- Once: import the source video to Sume, detach its audio, and transcribe it with sentence segments.
- Once, outside Sume: translate the sentences into each language. Sume has no translation endpoint in its public reference, so this step is yours.
- Per language: one TTS job from the translated script, with the language code and a voice for that language.
- Per language: join the sentence slices with Timeline audio if you split them, then one Timeline render over the video.
- Optional per language: a caption job, which is a new job, not a restyle of the source.
What does the fan-out look like?
This script submits every language's TTS job first, then collects the results, so the jobs run in parallel instead of one after another. It uses a fixed idempotency key per language, so running it twice does not bill twice. Voice ids come from your workspace, one per language.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
LINES = {"es": ("Hola, esta es nuestra nueva línea.", os.environ["VOICE_ES"]),
"de": ("Hallo, das ist unsere neue Linie.", os.environ["VOICE_DE"])}
FMT = {"container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le"}
def submit(lang, text, voice):
body = {"transcript": text, "language": lang, "voice": {"id": voice},
"mode": "async", "output_format": FMT}
r = requests.post(API + "/v1/tts-1.0/generate", json=body, timeout=60,
headers={**H, "Idempotency-Key": f"short-dub-{lang}-v1"})
r.raise_for_status()
return r.json()["data"]["result_url"]
def find(obj, key):
if isinstance(obj, dict) and key in obj:
return obj[key]
for v in (obj.values() if isinstance(obj, dict) else obj if isinstance(obj, list) else []):
if find(v, key) is not None:
return find(v, key)
urls = {lang: submit(lang, t, v) for lang, (t, v) in LINES.items()}
for lang, url in urls.items():
for _ in range(60):
r = requests.get(url, headers=H, timeout=60)
if r.status_code != 409:
break
time.sleep(3)
r.raise_for_status()
print(lang, find(r.json(), "audio_url"), find(r.json(), "duration_seconds"))How many jobs can you queue at once?
Jobs queue inside your workspace's plan limits. A submit response carries generation_limits, including the plan's concurrency, its queued-job limit, the remaining queue capacity and a wave_size_hint. Read that after the first submit rather than hard-coding eight; if you pass the limit the API returns a queue capacity rejection, so submit in waves. The generation admission page has the rules.
What does it cost?
| Step | Rate | Count | Subtotal |
|---|---|---|---|
| Audio detach | $0.01 per job | 1 | $0.01 |
| Speech-to-text | $0.01 per audio minute | 1 minute | $0.01 |
| TTS (assumed 900 characters per language) | $0.0475 per 1,000 characters | 8 | $0.342 |
| Timeline render | $0.10 per output minute, rounded up | 8 | $0.80 |
| Total, no captions | about $1.16 | ||
| Captions (optional) | $0.20 per video up to 60 seconds | 8 | $1.60 |
| Total with captions | about $2.76 |
Where do the eight results go?
Each TTS job returns an audio URL and a duration, which is what the script above prints. The duration is the number to compare against the original clip: if a translated line runs long, shorten the translation rather than speed the voice up past what sounds natural. TTS 1.0 does accept generation_config.speed between 0.6 and 1.5, but a small change goes further than the extremes.
From there each language is one Timeline render: the original video, with its own audio replaced by that language's track. Keep the source file and the per-language audio side by side, named by language code, so a re-run of one market never touches the others.
What if one language fails?
A failed job for one language does not stop the others, because each is its own job. Check the status of each result URL, resubmit only the failed language, and keep the idempotency key you used for the first try unless you changed the script. A changed script needs a new key, since the same key with a different body is a different request.
One practical order of work: run a single language end to end first, watch the usage block on each response, and only then submit the other seven. That way a mistake in the script format costs one job, not eight, and the total you computed above is checked against a real response before the batch.
What the total leaves out
For a feature-level walk through the same chain, see Translate a video's voiceover by API. Change the 900-character assumption and the TTS line moves; the render line does not, because it is priced by output minute.
- Your translation step, whatever model or person does it.
- The import step, which this table does not price; confirm it in
GET /v1/catalog. - Retries: resend with the same idempotency key to avoid a second job; a changed script is a new job and a new charge.
- Lip sync. Re-syncing a face to a new language is a separate surface with its own limits, covered in lip sync vs dubbing.
Sources
Related posts
More in Developers
- FLUX 3 bounding box to a mask_url: Python region edit on Sume
FLUX 3 Image boxes use [top, left, bottom, right] on a 0-1000 grid. Convert one to an RGBA mask with Pillow and run the region edit on Sume's GPT Image 2.5.
- FLUX 3 Image on OpenRouter: n=1, seed, base64 vs Sume
OpenRouter lists FLUX.3 Image with one image per call, a seed and base64 PNG output. How each differs from Sume's POST /v1/images, where FLUX 3 is not listed.
- FLUX 3 Image on Replicate: safety_tolerance 0-4 vs Sume
Replicate's FLUX 3 Image form has safety_tolerance 0-4, grounding, output_quality and 768sq-4k. Which of those inputs Sume's image API has, and what it returns.
- GPT Image 2.5 curl command: generate and download in a shell
A copy-paste curl call to Sume's POST /v1/images for GPT Image 2.5, with jq to pull the URL, download the file, and a check for the 202 job response.
Written by Sume