Turn Sume STT word timings into an SRT file in 30 lines of Python
Sume speech-to-text returns words[] with start and end seconds. Group them into SRT cues of 42 characters or 5 seconds, and see what the caption endpoint takes.

Sume's POST /v1/stt-1.0/transcribe always returns words[], each { word, start, end } in seconds from the start of the audio, so an SRT file is a grouping job: collect words into cues, break on sentence ends, a character budget and a time budget, then format the timestamps as HH:MM:SS,mmm. The function below does that with no dependencies. It takes the words array from a finished job and returns the SRT text.
Why you group words yourself
The API reference describes the STT result as text plus word timings, and a request can also ask for segmentation: { mode: "sentence" } to get sentence segments derived from the same words. Sentences are often too long for a subtitle line, though. A long sentence of 25 words needs to be split into two or three cues, and only you know your reading-speed budget. That is why the code below applies its own limits.
The function
Tune max_chars (42 is a common single-line limit) and max_seconds. The cue breaks early when a word ends in a period, question mark or exclamation mark.
def fmt(t):
ms = round(t * 1000)
h, ms = divmod(ms, 3_600_000)
m, ms = divmod(ms, 60_000)
s, ms = divmod(ms, 1000)
return f"{h:02}:{m:02}:{s:02},{ms:03}"
def words_to_srt(words, max_chars=42, max_seconds=5.0):
cues, cur = [], []
def flush():
if cur:
cues.append((cur[0]["start"], cur[-1]["end"], " ".join(w["word"] for w in cur)))
cur.clear()
for w in words:
text = " ".join(x["word"] for x in cur + [w])
if cur and (len(text) > max_chars or w["end"] - cur[0]["start"] > max_seconds):
flush()
cur.append(w)
if w["word"].endswith((".", "?", "!")):
flush()
flush()
return "\n".join(
f"{i}\n{fmt(a)} --> {fmt(b)}\n{t}\n" for i, (a, b, t) in enumerate(cues, 1)
)
Using it with a job result
Poll the job as described in jobs and results, read words from the result, and call words_to_srt(result["words"]). Write the string to a .srt file with UTF-8 encoding. Check the first and last cue against the audio before you ship a long file.
What Sume's caption endpoint takes instead
If your goal is subtitles burned into the picture, you do not need an SRT. The video captions endpoint transcribes the clip itself, or takes authored cues or words with text, start and end. The docs are explicit that SRT uploads are unsupported there, so the SRT you build here is for players, platforms and editors that read the file, while burned-in captions go through the caption job.
| Source | What it gives | Note |
|---|---|---|
| Sume STT 1.0 | words[] with start and end, optional sentence segments | Always returned; no flag to switch off |
| OpenAI whisper-1 | Word and segment timestamps | Per OpenAI's guide, only whisper-1 supports timestamp granularities for captioning |
| MAI-Transcribe-2-Streaming | Real-time partials, first one just over 100 ms | Microsoft positions it for live dictation and subtitling |
| Sume video captions | Burned-in captions from speech or authored cues | $0.20 per job for clips up to 60 seconds |
Limits worth knowing
The function does not translate, does not fix names the model misheard, and does not place two speakers on separate lines. Sume fixes diarization and audio-event tagging on the server side, so the words come back as one stream. If you need speaker labels, that is a different tool. For anything longer than 10 minutes, split the audio first and shift the times, as in this chunking walkthrough.
Sources
Related posts
More in Developers
- Count Sume job replays with idempotency_hit after a client retry
A Sume job submit returns data.idempotency_hit. Log it after a retried video submit to prove the second call replayed the first job, not a new one.
- Sume submit budgets: 120 to 1,200 writes a minute for an Omni batch
Sume rate limits submits per plan: Free 120, Pro 300, Startup 600 and Scale 1,200 a minute; reads get 40 times that. Why queue size, not the rate, paces Omni.
- Sume sync mode waits at most 30 seconds: short TTS vs long scripts
mode sync is a bounded wait, clamped to 0-30 seconds, not a promise the audio is ready. When to use it, and when to go async or webhook.
- Sume TTS 1.0 rejects model and model_id: use the router to pick
TTS 1.0 has no engine picker and returns 400 for model or model_id. The TTS router takes a required model from its catalog. Compare with ElevenLabs model tiers.
Written by Sume