Product clip voiceover in 5 languages: 20 seconds, $0.67 on Sume

A silent 20-second product clip voiced in five languages costs up to $0.67 on Sume: about $0.13 each for 300 characters of TTS, a concat and a render.

5 min readSume
All posts

Voicing a silent 20-second product clip in five languages costs up to $0.62 on Sume, which is $0.124 per language for 300 characters of TTS, a concat to read the length, and a one-minute timeline render. Bilibili's listed free Index-Translate API translates the script.

Assumptions: 300 characters per language (15 characters per second) and a spine of 60 seconds or less. Rates are from the pricing page and Timeline docs, read 2026-10-05.

Per language and for the set

The render is a one-minute minimum, so a 20-second clip pays for a minute.

Cost of a 20-second voiced clip by language count, up to, at list rates (read 2026-10-05)
LanguagesTTSConcatsRendersTotal, up to
1$0.0143$0.01$0.10$0.124
3$0.0428$0.03$0.30$0.373
5$0.0713$0.05$0.50$0.621
10$0.1425$0.10$1.00$1.242

Batch by clip length

Since a render is billed per output minute, a clip under a minute costs the same as one at 60 seconds. If you have several short clips, put them in one render with several video[] slots and one longer spine rather than rendering each separately.

Render one language

Set VIDEO_URL to the silent clip and SPEECH_URL to the TTS artifact for one language.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def run(path, body, key):
    r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
    r.raise_for_status()
    job = r.json()["data"]
    while not requests.get(job["status_url"], headers=H, timeout=30).json()["data"]["terminal"]:
        time.sleep(2)
    res = requests.get(job["result_url"], headers=H, timeout=30)
    res.raise_for_status()
    return res.json()["data"]["result"]

VIDEO, SPEECH = os.environ["VIDEO_URL"], os.environ["SPEECH_URL"]   # both on media.sume.com
joined = run("/v1/timeline-1.0/audio", {"operation": "concat", "parts": [{"url": SPEECH}]}, "dub-0412-de-spine")
secs = joined["duration_seconds"]
body = {"audio": {"url": joined["audio_url"], "duration_seconds": secs},
        "video": [{"source_url": VIDEO, "start": 0, "duration": secs}]}
out = run("/v1/timeline-1.0/render", body, "dub-0412-de-render")
print(out["video_url"])

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume