Translate an SRT and burn it in: Sume caption cues, limits, Python

Sume takes no SRT upload, but caption cues take the same text and times. A Python converter, the 200-cue and 60-second limits, and which fonts apply.

5 min readSume
All posts

Once your SRT is translated, convert each block to a cue with text, start and end in seconds, and send the list as cues on a caption job. Sume does not accept an SRT upload, and says so in its docs, but a cue carries the same three things an SRT block does, and a cue job skips speech-to-text entirely.

This follows the video captions docs, read 2026-10-03. The translation itself happens outside Sume, since the public API has no translation call.

What are the limits of a cue job?

A longer video needs its SRT cut into chunks of under 60 seconds against matching clips, which is a trim step on your side before captioning. Sume has video trim for that.

  • Up to 200 cues per job.
  • Each cue's text is at most 400 characters.
  • Cue start and end times stay within 60 seconds, so the video should be 60 seconds or less.
  • The job costs $0.20 per video up to 60 seconds.
  • cues, segments, words and script_text are mutually exclusive on one request.

A converter from SRT to cues

This script reads a translated SRT file, converts the times to seconds, trims each text to 400 characters, and writes cues.json ready to paste into a request. It needs only Python's standard library.

import re, json, sys

def ts(s):
    h, m, rest = s.strip().split(":")
    sec, ms = rest.replace(",", ".").split(".")
    return int(h) * 3600 + int(m) * 60 + int(sec) + int(ms) / 1000

def srt_to_cues(text):
    cues = []
    for block in re.split(r"\n\s*\n", text.strip()):
        lines = block.strip().splitlines()
        if len(lines) < 3 or "-->" not in lines[1]:
            continue
        a, b = lines[1].split("-->")
        cues.append({"text": " ".join(lines[2:])[:400],
                     "start": round(ts(a), 2), "end": round(ts(b), 2)})
    return cues

cues = srt_to_cues(open(sys.argv[1], encoding="utf-8").read())
print(len(cues), "cues; last ends at", cues[-1]["end"] if cues else 0)
json.dump({"cues": cues}, open("cues.json", "w"), ensure_ascii=False)

Which fonts and languages can you burn?

What caption styles draw, read 2026-10-03
Script of the textDocumented style choiceCaveat
Latin letters (Spanish, French, German)slam, punch, tiktok-greenLatin display faces
Koreanblack-outline and the other Hangul identitiesLatin styles return 400 for Hangul text
Japanese, Chinese, ArabicNone documentedTest a clip before a batch
Accented Latin textSame as LatinCheck diacritics on the style you pick

What do you lose against an SRT file?

An SRT is a sidecar: a viewer can switch it on or off, and the platform can show it in several languages. A burned-in caption is part of the picture, so there is one language per output video. If you need several languages on one upload, an SRT per language on a platform that accepts them is the right tool, and burning is for the cases where the platform shows none, or where you want the style.

Sume's language field on a caption job is only a hint for speech-to-text, and it does not pick the style or the font. With cues there is no speech-to-text at all, so the field has nothing to do.

What if the video is longer than 60 seconds?

Cue times are limited to 60 seconds, which fits a Short or an ad but not a tutorial. For a longer video, cut it into clips of at most 60 seconds, shift each chunk's cue times so they start at zero, and run one cue job per clip. Each job is its own $0.20 charge, so a 5-minute video in five clips is five jobs.

Doing the shifting in code is easy: subtract the clip's start time from each cue in that window and drop cues that fall outside it. Join the captioned clips back together afterwards in a Timeline render, with the original audio as the spine. Whether that is worth it, against uploading an SRT sidecar to a platform that accepts one, depends on whether you need the caption to be part of the picture.

Check the clip boundaries by eye on the first batch. A caption that appears for a fraction of a second at the start of a chunk usually means a cue was shifted by the wrong offset, and fixing the offset in code is better than editing cues one by one. Keep the original SRT untouched as the source of truth and generate every chunk's cues from it.

A cue spanning a cut is the one edge case: split it at the boundary into two cues, one in each clip, with the same text or a sensible break, instead of letting it overrun.

Before you send a batch

Send one video and one language first and look at the result on a phone-sized screen; line breaks in a translated sentence are the most common surprise, as translations run longer than the source. Keep the translated SRT and the cues.json it produced together, so a fix to one line is a one-line change and a re-render. The step-by-step for a plain English file is in burn an SRT file into a video.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume