Localize on-screen text per language: Python caption fan-out

Burn one set of translated cues per language onto a clean clip with Sume: one job each, its own key, asyncio fan-out, $0.20 a job. A runnable script.

5 min readSume
All posts

To put on-screen text in several languages on one clip, submit one POST /v1/video-captions job per language with that language's cues, each with its own Idempotency-Key, then poll each job to a terminal status. Each accepted job reserves $0.20 for a video up to 60 seconds, so six languages cost $1.20. The script below runs the jobs concurrently with asyncio and prints one result per language.

All endpoints and fields come from Video captions and Jobs and results, and the translations themselves are yours: Sume's docs describe no translation route.

What does one job per language look like?

cues (or segments) are phrase-level overlay cards with text, start, and end in seconds. Passing them skips speech-to-text, which also means the clip may be silent. script_text, words, cues, and segments are mutually exclusive, so send only one.

video_url must be a fetchable public HTTPS URL. Keep the source clip clean (no burned text) so each language starts from the same pixels.

import asyncio, json, os, urllib.request
API, KEY = "https://api.sume.com", os.environ["SUME_API_KEY"]
VIDEO = "https://media.sume.com/artifacts/artf_demo/clean.mp4"
CUES = {
    "es": [{"text": "Envio gratis hoy", "start": 0.5, "end": 3.0}],
    "de": [{"text": "Heute versandkostenfrei", "start": 0.5, "end": 3.0}],
}
def call(method, path, body=None, idem=None):
    h = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
    if idem:
        h["Idempotency-Key"] = idem
    data = json.dumps(body).encode() if body else None
    req = urllib.request.Request(API + path, data, h, method=method)
    with urllib.request.urlopen(req) as r:
        return json.load(r)
async def caption(lang, cues):
    body = {"video_url": VIDEO, "cues": cues, "mode": "async"}
    job = await asyncio.to_thread(call, "POST", "/v1/video-captions", body, f"promo-{lang}-001")
    while True:
        s = await asyncio.to_thread(call, "GET", f"/v1/jobs/{job['request_id']}/status")
        if s["terminal"]:
            break
        await asyncio.sleep(s.get("next_poll_after_seconds") or 5)
    return lang, s["sume_status"], job["request_id"]
async def main():
    print(await asyncio.gather(*(caption(k, v) for k, v in CUES.items())))
asyncio.run(main())

Why a different Idempotency-Key per language?

The jobs page says to reuse a key only for the same operation and payload, such as retrying after a client timeout. A Spanish job and a German job are different payloads, so they need different keys. The script builds the key from the language and a run suffix; change the suffix when you deliberately want a new job with the same cues.

How does polling work?

A submit with mode: "async" returns a job envelope whose request_id is the job id. Poll GET /v1/jobs/{id}/status until terminal is true, and honor next_poll_after_seconds when present. sume_status is one of queued, processing, completed, failed, or canceled.

Do not resubmit the paid request because your process timed out; the job keeps running and billing. Read the finished video from GET /v1/jobs/{id}/result, which answers 409 job_not_completed until the job is complete.

Job statuses, from docs.sume.com Jobs and results, read 2026-10-02.
sume_statusTerminal
queuedNo
processingNo
completedYes
failedYes
canceledYes

What should you check before running six languages?

Fonts first. The docs describe Latin styles and Hangul styles; Korean copy sent to slam, punch, or tiktok-green is rejected with 400 (caption_hangul_text_latin_style). Pick a Hangul style for Korean cues, such as black-outline, which is also what an omitted style resolves to for Korean. For other scripts the docs list nothing, so test one clip.

Then length. Translated text rarely matches the original's length, so check that each cue fits its time window. You can tune design.phrasing (max_words, max_chars) and design.typography.safe_width_ratio per request; out-of-range numbers return a 400 at request time instead of billing a bad render.

If you would rather not poll, send mode: "webhook" with a public HTTPS webhook_url and keep polling as a fallback for missed deliveries, as the webhooks page describes.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume