Turn Sume STT word timings into an SRT file in 30 lines of Python

Sume speech-to-text returns words[] with start and end seconds. Group them into SRT cues of 42 characters or 5 seconds, and see what the caption endpoint takes.

6 min readSume
All posts

Sume's POST /v1/stt-1.0/transcribe always returns words[], each { word, start, end } in seconds from the start of the audio, so an SRT file is a grouping job: collect words into cues, break on sentence ends, a character budget and a time budget, then format the timestamps as HH:MM:SS,mmm. The function below does that with no dependencies. It takes the words array from a finished job and returns the SRT text.

Why you group words yourself

The API reference describes the STT result as text plus word timings, and a request can also ask for segmentation: { mode: "sentence" } to get sentence segments derived from the same words. Sentences are often too long for a subtitle line, though. A long sentence of 25 words needs to be split into two or three cues, and only you know your reading-speed budget. That is why the code below applies its own limits.

The function

Tune max_chars (42 is a common single-line limit) and max_seconds. The cue breaks early when a word ends in a period, question mark or exclamation mark.

def fmt(t):
    ms = round(t * 1000)
    h, ms = divmod(ms, 3_600_000)
    m, ms = divmod(ms, 60_000)
    s, ms = divmod(ms, 1000)
    return f"{h:02}:{m:02}:{s:02},{ms:03}"


def words_to_srt(words, max_chars=42, max_seconds=5.0):
    cues, cur = [], []

    def flush():
        if cur:
            cues.append((cur[0]["start"], cur[-1]["end"], " ".join(w["word"] for w in cur)))
            cur.clear()

    for w in words:
        text = " ".join(x["word"] for x in cur + [w])
        if cur and (len(text) > max_chars or w["end"] - cur[0]["start"] > max_seconds):
            flush()
        cur.append(w)
        if w["word"].endswith((".", "?", "!")):
            flush()
    flush()
    return "\n".join(
        f"{i}\n{fmt(a)} --> {fmt(b)}\n{t}\n" for i, (a, b, t) in enumerate(cues, 1)
    )

Using it with a job result

Poll the job as described in jobs and results, read words from the result, and call words_to_srt(result["words"]). Write the string to a .srt file with UTF-8 encoding. Check the first and last cue against the audio before you ship a long file.

What Sume's caption endpoint takes instead

If your goal is subtitles burned into the picture, you do not need an SRT. The video captions endpoint transcribes the clip itself, or takes authored cues or words with text, start and end. The docs are explicit that SRT uploads are unsupported there, so the SRT you build here is for players, platforms and editors that read the file, while burned-in captions go through the caption job.

Where subtitle timings can come from (vendor facts read 2026-10-04)
SourceWhat it givesNote
Sume STT 1.0words[] with start and end, optional sentence segmentsAlways returned; no flag to switch off
OpenAI whisper-1Word and segment timestampsPer OpenAI's guide, only whisper-1 supports timestamp granularities for captioning
MAI-Transcribe-2-StreamingReal-time partials, first one just over 100 msMicrosoft positions it for live dictation and subtitling
Sume video captionsBurned-in captions from speech or authored cues$0.20 per job for clips up to 60 seconds

Limits worth knowing

The function does not translate, does not fix names the model misheard, and does not place two speakers on separate lines. Sume fixes diarization and audio-event tagging on the server side, so the words come back as one stream. If you need speaker labels, that is a different tool. For anything longer than 10 minutes, split the audio first and shift the times, as in this chunking walkthrough.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume