Translate all subtitle cues in one JSON request, then validate

Index-Translate lists JSON format preservation. Send a clip's cues as one JSON array, then validate count, timings and length before Sume burns them.

5 min readSume
All posts

You can translate every cue of a clip in a single request by sending the cues as a JSON array and asking the model to keep the structure, then validating the reply before Sume burns it. Index-Translate's README lists format preservation for JSON, CSV and code among its instruction-following features, which is exactly what a cue list needs: the text changes, the timings and keys do not.

One request per clip keeps the context together, so a sentence split across two cues is translated as a whole thought. The cost is that a malformed reply spoils the batch, so the validation step is not optional.

Why send the cues together?

Cue-by-cue translation is simple, but it loses context. A cue that reads "it costs" on its own gives the model nothing to work with, and a pronoun in the next line can flip the meaning. A JSON array of up to 200 cues, the most Sume accepts in one caption request, lets the model see the neighbours.

The model card for the FP8 checkpoint sets a maximum model length of 4096 tokens. A 60-second clip fits comfortably, but anything longer should be sent in windows of cues that stay inside it.

Per-cue versus whole-array translation, read 2026-10-04
ApproachContextFailure modeRequests per clip
One cue per requestNone across linesPronouns and fragments mistranslatedOne per cue
Whole JSON arrayNeighbouring lines visibleMalformed or reordered JSONOne

What do you validate?

The check below compares the model's reply with the source cues: same count, same start and end, non-empty text under 400 characters. Those are Sume's own cue rules, along with end greater than start and times inside 60 seconds, so a reply that passes can go straight into a caption request.

import json

def validate(source, reply_text):
    try:
        out = json.loads(reply_text)
    except json.JSONDecodeError as err:
        return [f"not JSON: {err}"]
    if len(out) != len(source):
        return [f"expected {len(source)} cues, got {len(out)}"]
    problems = []
    for i, (a, b) in enumerate(zip(source, out)):
        if (a["start"], a["end"]) != (b.get("start"), b.get("end")):
            problems.append(f"cue {i}: timing changed")
        text = str(b.get("text", "")).strip()
        if not text or len(text) > 400:
            problems.append(f"cue {i}: text empty or over 400 chars")
    return problems

source = [{"text": "你好", "start": 0.0, "end": 1.2}]
reply = '[{"text": "Hello", "start": 0.0, "end": 1.2}]'
print(validate(source, reply))

What do you do when validation fails?

Once the cues validate, send them as cues to video captions with a public HTTPS video_url. A standalone job is priced at $0.20 for videos up to 60 seconds under the current estimate, so every failed render you avoid is real money saved. The request itself is shown in the cue-translation walkthrough.

  • Retry the whole array once with a shorter instruction that repeats the structure rule.
  • If it fails again, split the cues in half and translate each half.
  • As a last resort, fall back to one cue per request for the failed range.
  • Never edit timings by hand to make a reply pass; fix the source of the mismatch.

Should you also check the language of the reply?

Yes. A model asked to translate to English can leave a cue untouched if it judges the text already fine, or mix two languages in one line. A cheap check is to confirm that the translated text differs from the source for every cue where it should, and that no cue still contains the source script when your target is Latin. Sume's Latin styles reject Korean text with a 400, so a Hangul leftover in an English cue would fail the render and cost you a retry.

Run these checks in the same function as the structure check, so one pass decides whether a window is safe to post.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume