Convert an SRT file to Sume caption cues in Python
Sume captions take no SRT upload, but cues carry the same start, end and text. A 26-line Python script turns an SRT into cues and posts them for $0.20.

Sume's video caption endpoint does not take an SRT file. The docs say it does not support SRT uploads, and that the way to send phrase-level text is cues (or its alias segments), each with text, start and end in seconds. An SRT block is already that shape, so conversion is a short parse: read the time range, read the text lines, and post the list. The script below does it in 26 lines of standard-library Python and runs when you pass it an SRT path and a public HTTPS video URL.
What an SRT block maps to
An SRT entry has an index, a start --> end line with HH:MM:SS,mmm times, and one or more text lines. A cue has text, start and end. The index drops out, the comma in the milliseconds becomes a decimal point, and a two-line subtitle becomes one cue with the lines joined by \n, which the docs allow.
| Item | SRT | Sume cue |
|---|---|---|
| Text | One or more lines | Up to 400 characters, lines joined with a newline |
| Start and end | HH:MM:SS,mmm | Seconds as numbers, end must be greater than start |
| Count | Unlimited | 1 to 200 cues per request |
| Timeline | Any length | Each time at most 60 seconds |
| Speech-to-text | Not applicable | Not run when cues are sent |
The script
Run it as python srt.py captions.srt https://media.sume.com/artifacts/artf_demo/clean.mp4. It refuses to send more than 200 cues or anything past 60 seconds, so you fix the file locally before you spend $0.20 on a failed request. The idempotency key is built from the filename, so a retry of the same file does not queue a second job; change the key if you edit the file, because the same key with a different body is a conflict.
import json, os, re, sys, urllib.request
def secs(t):
h, m, rest = t.strip().split(":")
s, ms = rest.split(",")
return int(h) * 3600 + int(m) * 60 + int(s) + int(ms) / 1000
cues = []
for block in re.split(r"\n\s*\n", open(sys.argv[1], encoding="utf-8").read().strip()):
lines = block.splitlines()
start, end = (secs(x) for x in lines[1].split("-->"))
cues.append({"text": "\n".join(lines[2:])[:400], "start": start, "end": end})
assert 1 <= len(cues) <= 200 and cues[-1]["end"] <= 60, "cues must fit 200 and 60 s"
req = urllib.request.Request(
"https://api.sume.com/v1/video-captions",
data=json.dumps({"video_url": sys.argv[2], "cues": cues}).encode(),
headers={
"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": "srt-cues-" + os.path.basename(sys.argv[1]),
},
)
print(json.loads(urllib.request.urlopen(req).read()))Before you send it
Check three things. First, the video must be a public HTTPS URL that Sume can fetch; localhost, private networks and signed or private URLs are rejected. Second, the video must be 60 seconds or shorter, since the standalone job price is $0.20 for up to 60 seconds. Third, the cues should not overlap each other unless you want two cards visible at once.
If your SRT came from an editor that adds markup such as <i> tags or {\an8} position codes, strip it before the join. The docs do not describe tag handling, so the safe assumption is that characters you send are the characters that get drawn.
When STT is better than an SRT
If you do not already have subtitles, skip the SRT entirely. Send only video_url and Sume transcribes and burns in one job. If you have a script but not timings, send script_text and let speech-to-text supply the times. Cues are the right tool when you own both wording and timing, for example when the subtitles were translated and reviewed by a person.
For a longer video, split it into 60 second pieces first. The $0.20 price covers videos of up to 60 seconds, so a three minute video is three jobs and $0.60, and each piece needs its own SRT with times that start at zero. Sume does not accept ffmpeg commands, so cut with POST /v1/video-trim and shift the cue times by the piece's start before you post them.
Common conversion errors
A byte-order mark at the start of the file breaks the first index line in some parsers; open the file as utf-8-sig if you see an unexpected first character. Windows line endings can leave stray carriage returns in cue text, which the splitlines call above already handles. Overlapping or zero-length cues fail validation, because the request schema requires end to be greater than start.
If you hit the 400 character limit on a cue, you probably have a subtitle that holds a full paragraph. Split it at a sentence boundary into two cues, dividing its time range in proportion to the text length.
Sources
Related posts
More in Developers
- Cost per ad variant: build a ledger from usage.cost on Sume
Every completed /v1/videos poll carries usage.cost. Sum it by hook and ending to get the cost per ad variant before media spend. Node script and the caveats.
- Create an AI avatar and its first talking video in one bash script
Two Sume jobs in order: create the avatar, wait, then render a talking video with its handle. A bash script with curl and jq, plus the cost of both steps.
- CTA end card for an AI video ad: use last_frame on Sume
End an ad clip on your CTA card by sending it as a last_frame image on /v1/videos. Which models accept it, which do not, and the Timeline alternative.
- How do I narrate a DIY tutorial step by step with a TTS API?
Narrate an 8-step DIY tutorial with one TTS job per step: 1,570 characters, $0.10 on Sume. Why per-step jobs make a fixed step a 1-cent redo.
Written by Sume