Make an SRT file from Sume STT word timings in Python (7 words a cue)

Sume STT always returns words[] with word, start and end in seconds. This Python function turns the list into an SRT file, seven words a cue.

3 min readSume
All posts

Sume STT returns word-level timings with every transcription: a words list of objects with word, start and end, in seconds from the start of the audio. There is no flag to turn them on. That is enough to write an SRT subtitle file yourself, as in the function below, which groups seven words into each cue.

The request and result fields are from Sume's OpenAPI description of POST /v1/stt-1.0/transcribe, read on 2026-10-09. STT is priced at $0.01 per audio minute.

The function

SRT wants a number, a time range in HH:MM:SS,mmm form, the text and a blank line. The function below formats the times from seconds and ends each cue on the end time of its last word. It uses a sample words list so you can run it as-is; in a real script you would pass the words from your STT job result.

def ts(sec):
    ms = int(round(sec * 1000))
    h, ms = divmod(ms, 3600000)
    m, ms = divmod(ms, 60000)
    s, ms = divmod(ms, 1000)
    return f"{h:02}:{m:02}:{s:02},{ms:03}"

def words_to_srt(words, per_cue=7):
    cues = []
    for i in range(0, len(words), per_cue):
        chunk = words[i:i + per_cue]
        text = " ".join(w["word"] for w in chunk)
        cues.append(f"{len(cues) + 1}\n{ts(chunk[0]['start'])} --> "
                    f"{ts(chunk[-1]['end'])}\n{text}\n")
    return "\n".join(cues)

sample = [{"word": w, "start": i * 0.4, "end": i * 0.4 + 0.35}
          for i, w in enumerate("this is a short test of the srt writer".split())]
print(words_to_srt(sample))

Where this stops being enough

Fixed word counts split sentences in odd places. If you need cues that follow sentences, ask STT for segmentation with mode sentence, which groups words into sentences on terminal punctuation and splits unpunctuated runs on silence, then build one cue per segment. The boundary_lead_ms setting, from 0 to 500 with a default of 70, controls how much time is carried past a sentence's last word.

Also remember that SRT carries no styling. A sidecar file lets a player show or hide subtitles, but it does not burn them into the picture.

SRT or burned-in captions

If the platform you post to reads sidecar files, SRT is free beyond the STT charge. For a short vertical video on a social feed, where subtitles must be in the picture, Sume's caption job burns them in for $0.20 per job on clips up to 60 seconds, with styles such as slam, punch and tiktok-green. That job runs its own speech recognition unless you supply words, cues or script_text, so you do not need to pass it the SRT.

  • Sidecar file for players that support it: this function, no extra cost.
  • Burned-in look for social video: the caption job.
  • Check the first cue's start time; very quiet openings can shift it.

Testing the output

Save the string to a file with the .srt extension and open it in a video player that supports subtitles. Check the first and last cue against the audio, and look for cues that are too long to read at a comfortable pace. A common rule of thumb is no more than two lines on screen at once.

If the audio has a long silent opening, the first word's start time will be late, which is correct. If you see the first cue starting at zero when the speech starts later, you are probably passing the wrong audio file, or a slice that was trimmed after the STT job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume