Index-Translate subtitles: translate cues, then burn them in

Bilibili's Index-Translate covers 150 languages with a free OpenAI-compatible API. Translate each cue, then burn the result with Sume video captions.

5 min readSume
All posts

To get translated, burned-in subtitles from a video, translate each caption cue with Index-Translate, then send the cues to Sume's video captions endpoint, which burns them in without running speech-to-text. Index-Translate is Bilibili's translation model family, released on September 30, 2026 under Apache-2.0 and listed for 150 languages. Its site offers a free OpenAI-compatible API, so the translation step costs nothing but a request per cue.

The two halves meet at one object: a cue with text, start and end in seconds. You keep the timings from your source subtitles, replace only the text, and Sume draws the result onto your clip.

What does Index-Translate give you?

The GitHub repository describes a family: text translation across 150 languages, plus speech models for subtitle and dubbing work. The text model's 35B-A3B preview is a mixture-of-experts model with 35B total and about 3B active parameters. The online demo page lists the API model id as Index-Translate-35B-A3B and calls the endpoint free and OpenAI-compatible. It states no rate limits, so treat the endpoint as a demo and keep volumes modest.

Index-Translate facts used in this post, read 2026-10-04
ItemWhat the vendor page says
Release dateSeptember 30, 2026
LicenceApache-2.0
Text languages150
APIFree public endpoint, OpenAI-compatible, model id Index-Translate-35B-A3B
Prompt on the model cardChinese instruction asking for the translation only, temperature 0
Rate limitsNot stated

What does Sume need from the cues?

Video captions takes a public HTTPS video_url and, instead of speech-to-text, accepts cues (or the alias segments) with text, start and end. In the Sume API schema each cue's text can be up to 400 characters, a request carries at most 200 cues, and start and end times stay inside 60 seconds. cues cannot be combined with script_text or words. A standalone caption job is priced at $0.20 for videos up to 60 seconds under the current fixed estimate; confirm it in GET /v1/catalog.

For Latin-script targets such as English, Spanish or French, name a Latin style like slam. Sume's Latin styles have no Hangul glyphs, so Korean copy needs a Hangul style instead.

How do you wire the two together?

The script below translates each cue and submits one caption job. It reads the Sume key from the environment and refuses to run without it. The Index-Translate page does not say whether its endpoint wants a key, so the script sends none; add an Authorization header if the endpoint asks for one.

import json, os, urllib.request

def post(url, body, headers):
    data = json.dumps(body).encode()
    req = urllib.request.Request(url, data, headers)
    with urllib.request.urlopen(req, timeout=120) as r:
        return json.load(r)

def translate(text, lang):
    prompt = f"请将以下文本翻译为{lang},直接输出翻译结果,不要进行任何解释。\n\n{text}"
    out = post("https://index-translate.bilibili.com/v1/chat/completions",
               {"model": "Index-Translate-35B-A3B", "temperature": 0,
                "messages": [{"role": "user", "content": prompt}]},
               {"Content-Type": "application/json"})
    return out["choices"][0]["message"]["content"].strip()

key = os.environ.get("SUME_API_KEY", "")
if not key:
    raise SystemExit("set SUME_API_KEY")
cues = [{"text": "你好,欢迎来到我们的店。", "start": 0.0, "end": 2.4}]
for c in cues:
    c["text"] = translate(c["text"], "English")
job = post("https://api.sume.com/v1/video-captions",
           {"video_url": "https://example.com/clip.mp4", "style": "slam", "cues": cues},
           {"Authorization": f"Bearer {key}", "Content-Type": "application/json",
            "Idempotency-Key": "index-translate-demo-001"})
print(job)

What should you check before you burn?

Translation changes length. A line that fit 2.4 seconds in Chinese can run longer in English, so check reading speed on the translated cues before paying for a render. The model card also notes greedy decoding at temperature 0, which is what the script sets. For the SRT-file route and the Sume limits in detail, see translate an SRT and burn it in.

Poll the job with GET /v1/jobs/{id}/status, then read the captioned video_url from the result.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume