Convert an SRT file to Sume caption cues in Python

Sume captions take no SRT upload, but cues carry the same start, end and text. A 26-line Python script turns an SRT into cues and posts them for $0.20.

5 min readSume
All posts

Sume's video caption endpoint does not take an SRT file. The docs say it does not support SRT uploads, and that the way to send phrase-level text is cues (or its alias segments), each with text, start and end in seconds. An SRT block is already that shape, so conversion is a short parse: read the time range, read the text lines, and post the list. The script below does it in 26 lines of standard-library Python and runs when you pass it an SRT path and a public HTTPS video URL.

What an SRT block maps to

An SRT entry has an index, a start --> end line with HH:MM:SS,mmm times, and one or more text lines. A cue has text, start and end. The index drops out, the comma in the milliseconds becomes a decimal point, and a two-line subtitle becomes one cue with the lines joined by \n, which the docs allow.

SRT fields and Sume cue limits, from the caption request schema read 2026-10-07
ItemSRTSume cue
TextOne or more linesUp to 400 characters, lines joined with a newline
Start and endHH:MM:SS,mmmSeconds as numbers, end must be greater than start
CountUnlimited1 to 200 cues per request
TimelineAny lengthEach time at most 60 seconds
Speech-to-textNot applicableNot run when cues are sent

The script

Run it as python srt.py captions.srt https://media.sume.com/artifacts/artf_demo/clean.mp4. It refuses to send more than 200 cues or anything past 60 seconds, so you fix the file locally before you spend $0.20 on a failed request. The idempotency key is built from the filename, so a retry of the same file does not queue a second job; change the key if you edit the file, because the same key with a different body is a conflict.

import json, os, re, sys, urllib.request


def secs(t):
    h, m, rest = t.strip().split(":")
    s, ms = rest.split(",")
    return int(h) * 3600 + int(m) * 60 + int(s) + int(ms) / 1000


cues = []
for block in re.split(r"\n\s*\n", open(sys.argv[1], encoding="utf-8").read().strip()):
    lines = block.splitlines()
    start, end = (secs(x) for x in lines[1].split("-->"))
    cues.append({"text": "\n".join(lines[2:])[:400], "start": start, "end": end})
assert 1 <= len(cues) <= 200 and cues[-1]["end"] <= 60, "cues must fit 200 and 60 s"

req = urllib.request.Request(
    "https://api.sume.com/v1/video-captions",
    data=json.dumps({"video_url": sys.argv[2], "cues": cues}).encode(),
    headers={
        "Authorization": "Bearer " + os.environ["SUME_API_KEY"],
        "Content-Type": "application/json",
        "Idempotency-Key": "srt-cues-" + os.path.basename(sys.argv[1]),
    },
)
print(json.loads(urllib.request.urlopen(req).read()))

Before you send it

Check three things. First, the video must be a public HTTPS URL that Sume can fetch; localhost, private networks and signed or private URLs are rejected. Second, the video must be 60 seconds or shorter, since the standalone job price is $0.20 for up to 60 seconds. Third, the cues should not overlap each other unless you want two cards visible at once.

If your SRT came from an editor that adds markup such as <i> tags or {\an8} position codes, strip it before the join. The docs do not describe tag handling, so the safe assumption is that characters you send are the characters that get drawn.

When STT is better than an SRT

If you do not already have subtitles, skip the SRT entirely. Send only video_url and Sume transcribes and burns in one job. If you have a script but not timings, send script_text and let speech-to-text supply the times. Cues are the right tool when you own both wording and timing, for example when the subtitles were translated and reviewed by a person.

For a longer video, split it into 60 second pieces first. The $0.20 price covers videos of up to 60 seconds, so a three minute video is three jobs and $0.60, and each piece needs its own SRT with times that start at zero. Sume does not accept ffmpeg commands, so cut with POST /v1/video-trim and shift the cue times by the piece's start before you post them.

Common conversion errors

A byte-order mark at the start of the file breaks the first index line in some parsers; open the file as utf-8-sig if you see an unexpected first character. Windows line endings can leave stray carriage returns in cue text, which the splitlines call above already handles. Overlapping or zero-length cues fail validation, because the request schema requires end to be greater than start.

If you hit the 400 character limit on a cue, you probably have a subtitle that holds a full paragraph. Split it at a sentence boundary into two cues, dividing its time range in proportion to the text length.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume