Course trailer in 4 languages from one video track: $0.57 on Sume

A 45-second course trailer in four languages reuses one video track: each language is about $0.14 of TTS, concat and a one-minute render, $0.57 in all.

5 min readSume
All posts

A 45-second course trailer in four languages costs up to $0.57 on Sume when you reuse one video track: $0.142 per language for the speech, a concat and a one-minute timeline render. The translation of the trailer script runs on Bilibili's listed free Index-Translate API.

Assumptions: 675 characters of speech per language (15 characters per second for 45 seconds) and spines of 60 seconds or less. Rates are from the pricing page and the Timeline docs, read 2026-10-05.

Per-language cost

The visual cuts are made once; only the audio spine changes.

Cost of the trailer by number of languages, up to, at list rates (read 2026-10-05)
LanguagesTTSConcatsRendersTotal, up to
1$0.032$0.01$0.10$0.14
2$0.064$0.02$0.20$0.28
4$0.128$0.04$0.40$0.57
10$0.321$0.10$1.00$1.42

Keep the cuts in sync

Different languages speak at different speeds, so a cut tied to a spoken line moves. Time your video slots from each language's own concat result, and keep on-screen text off the shot or burn it per language with caption cues.

Render one language

Run this once for each language's speech URL and trailer video.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def run(path, body, key):
    r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
    r.raise_for_status()
    job = r.json()["data"]
    while not requests.get(job["status_url"], headers=H, timeout=30).json()["data"]["terminal"]:
        time.sleep(2)
    res = requests.get(job["result_url"], headers=H, timeout=30)
    res.raise_for_status()
    return res.json()["data"]["result"]

VIDEO, SPEECH = os.environ["VIDEO_URL"], os.environ["SPEECH_URL"]   # both on media.sume.com
joined = run("/v1/timeline-1.0/audio", {"operation": "concat", "parts": [{"url": SPEECH}]}, "dub-0412-de-spine")
secs = joined["duration_seconds"]
body = {"audio": {"url": joined["audio_url"], "duration_seconds": secs},
        "video": [{"source_url": VIDEO, "start": 0, "duration": secs}]}
out = run("/v1/timeline-1.0/render", body, "dub-0412-de-render")
print(out["video_url"])

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume