One video in three languages: a timeline job per language

Keep the same video slots and swap the audio spine per language. A Timeline 1.0 render is $0.10 per output minute, so three 60 s versions cost $0.30.

4 min readSume
All posts

Render each language as its own Timeline 1.0 job: same video[] slots, a different Sume-hosted audio spine. The render is $0.10 per output minute with ffmpeg only, so three 60-second versions cost $0.30 plus the speech synthesis for each language.

What stays and what changes

The video slots, transitions and output size stay identical. Only audio.url and audio.duration_seconds change, because a translated voice is rarely the same length. Declared video[].start values are authoritative, so if a language runs longer you must lengthen the slots too.

Per-language render inputs (read 2026-10-03)
FieldSame across languages?Note
video[].source_urlYesSume-hosted clips
video[].start / durationNo, when speech length differsCoverage may trail the spine by at most 0.5 s
audio.urlNoYour TTS output for that language
audio.duration_secondsNo1 to 1800

Loop over languages

Every URL must already be on media.sume.com. The audio URLs below are placeholders for TTS results you already have; the slot timings are for a 24-second spine.

import os, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
VOICE = {
    "en": os.environ["VOICE_EN"],
    "es": os.environ["VOICE_ES"],
    "ko": os.environ["VOICE_KO"],
}
for lang, url in VOICE.items():
    body = {
        "audio": {"url": url, "duration_seconds": 24},
        "video": [{"source_url": os.environ["CLIP_URL"], "start": 0, "duration": 24}],
    }
    r = requests.post("https://api.sume.com/v1/timeline-1.0/render",
                      headers={**H, "Idempotency-Key": f"loc-render-{lang}-001"},
                      json=body, timeout=30)
    r.raise_for_status()
    print(lang, r.json()["status_url"])

Then captions

Burn per-language captions with cues in a separate video-captions job on each render, using a Hangul style for Korean. Lip movement still follows the source language; for on-camera speakers that is a separate lip-sync step.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume