Translate an SRT and burn it in: Sume caption cues, limits, Python
Sume takes no SRT upload, but caption cues take the same text and times. A Python converter, the 200-cue and 60-second limits, and which fonts apply.

Once your SRT is translated, convert each block to a cue with text, start and end in seconds, and send the list as cues on a caption job. Sume does not accept an SRT upload, and says so in its docs, but a cue carries the same three things an SRT block does, and a cue job skips speech-to-text entirely.
This follows the video captions docs, read 2026-10-03. The translation itself happens outside Sume, since the public API has no translation call.
What are the limits of a cue job?
A longer video needs its SRT cut into chunks of under 60 seconds against matching clips, which is a trim step on your side before captioning. Sume has video trim for that.
- Up to 200 cues per job.
- Each cue's text is at most 400 characters.
- Cue start and end times stay within 60 seconds, so the video should be 60 seconds or less.
- The job costs $0.20 per video up to 60 seconds.
cues,segments,wordsandscript_textare mutually exclusive on one request.
A converter from SRT to cues
This script reads a translated SRT file, converts the times to seconds, trims each text to 400 characters, and writes cues.json ready to paste into a request. It needs only Python's standard library.
import re, json, sys
def ts(s):
h, m, rest = s.strip().split(":")
sec, ms = rest.replace(",", ".").split(".")
return int(h) * 3600 + int(m) * 60 + int(sec) + int(ms) / 1000
def srt_to_cues(text):
cues = []
for block in re.split(r"\n\s*\n", text.strip()):
lines = block.strip().splitlines()
if len(lines) < 3 or "-->" not in lines[1]:
continue
a, b = lines[1].split("-->")
cues.append({"text": " ".join(lines[2:])[:400],
"start": round(ts(a), 2), "end": round(ts(b), 2)})
return cues
cues = srt_to_cues(open(sys.argv[1], encoding="utf-8").read())
print(len(cues), "cues; last ends at", cues[-1]["end"] if cues else 0)
json.dump({"cues": cues}, open("cues.json", "w"), ensure_ascii=False)Which fonts and languages can you burn?
| Script of the text | Documented style choice | Caveat |
|---|---|---|
| Latin letters (Spanish, French, German) | slam, punch, tiktok-green | Latin display faces |
| Korean | black-outline and the other Hangul identities | Latin styles return 400 for Hangul text |
| Japanese, Chinese, Arabic | None documented | Test a clip before a batch |
| Accented Latin text | Same as Latin | Check diacritics on the style you pick |
What do you lose against an SRT file?
An SRT is a sidecar: a viewer can switch it on or off, and the platform can show it in several languages. A burned-in caption is part of the picture, so there is one language per output video. If you need several languages on one upload, an SRT per language on a platform that accepts them is the right tool, and burning is for the cases where the platform shows none, or where you want the style.
Sume's language field on a caption job is only a hint for speech-to-text, and it does not pick the style or the font. With cues there is no speech-to-text at all, so the field has nothing to do.
What if the video is longer than 60 seconds?
Cue times are limited to 60 seconds, which fits a Short or an ad but not a tutorial. For a longer video, cut it into clips of at most 60 seconds, shift each chunk's cue times so they start at zero, and run one cue job per clip. Each job is its own $0.20 charge, so a 5-minute video in five clips is five jobs.
Doing the shifting in code is easy: subtract the clip's start time from each cue in that window and drop cues that fall outside it. Join the captioned clips back together afterwards in a Timeline render, with the original audio as the spine. Whether that is worth it, against uploading an SRT sidecar to a platform that accepts one, depends on whether you need the caption to be part of the picture.
Check the clip boundaries by eye on the first batch. A caption that appears for a fraction of a second at the start of a chunk usually means a cue was shifted by the wrong offset, and fixing the offset in code is better than editing cues one by one. Keep the original SRT untouched as the source of truth and generate every chunk's cues from it.
A cue spanning a cut is the one edge case: split it at the boundary into two cues, one in each clip, with the same text or a sensible break, instead of letting it overrun.
Before you send a batch
Send one video and one language first and look at the result on a phone-sized screen; line breaks in a translated sentence are the most common surprise, as translations run longer than the source. Keep the translated SRT and the cues.json it produced together, so a fix to one line is a one-line change and a re-render. The step-by-step for a plain English file is in burn an SRT file into a video.
Sources
Related posts
More in Developers
- Validate video duration and resolution in Python before you submit
Fetch GET /v1/videos/models and check duration, resolution and aspect_ratio per model in about 25 lines of Python, before a Sume video job fails.
- AI video API fallback: retry on another model when a job fails
Chain seedance-2.5, seedance-2 and kling-3 on Sume: poll status_url, read the job error category, and resubmit the brief to the next model.
- Virtual try-on API: which Sume call returns an image, which a video
Need a try-on photo or a try-on clip? On Sume the two catalog try-on Formats return video; a still comes from the image API. The table, plus one call for each.
- Test a webhook endpoint before go-live: a Sume CI gate (Python)
Use POST /v1/webhooks/test-deliveries to fire a signed webhook.test at your deployed URL and fail the deploy unless it answers 2xx. Python script included.
Written by Sume