Sume STT to an SRT file: build subtitles from sentence segments

Sume returns timed sentence segments, not an SRT. Turn them into a valid .srt file in Python for YouTube, Vimeo or a player, with the timestamp format.

5 min readSume
All posts

Sume speech-to-text does not output SRT or WebVTT. It returns the transcript text, word timings and, if you ask for segmentation: { mode: "sentence" }, gapless sentence segments with index, text, start and end in seconds. Writing an .srt from those segments is a loop and one timestamp function, shown below. Upload the file to YouTube, Vimeo or a player when you want selectable captions rather than burned-in ones.

Transcription is getting cheaper and more live: Microsoft's MAI-Transcribe-2-Streaming lists 60 languages at $0.54 an hour as an intro price through 2026 (read 2026-10-06 on the October tracker). Sume's batch rate is $0.01 per audio minute, or $0.60 an hour (pricing). The output is the difference: you get timings you can turn into whatever file format a platform wants.

Segments to SRT

import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]

def call(url, body=None, key=None):
    h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
    if key: h["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body else None
    return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]

def stamp(t):
    ms = round(t * 1000)
    h, ms = divmod(ms, 3600000)
    m, ms = divmod(ms, 60000)
    s, ms = divmod(ms, 1000)
    return f"{h:02}:{m:02}:{s:02},{ms:03}"

job = call("https://api.sume.com/v1/stt-1.0/transcribe", {
    "audio_url": os.environ["AUDIO_URL"], "language_code": "en", "duration_seconds": 300,
    "segmentation": {"mode": "sentence"}}, "srt-demo-v1")
while not job["terminal"]:
    time.sleep(job.get("next_poll_after_seconds") or 3)
    job = call(job["status_url"])
segments = call(job["result_url"])["result"]["segments"]
with open("episode.srt", "w", encoding="utf-8") as f:
    for s in segments:
        f.write(f"{s['index'] + 1}\n{stamp(s['start'])} --> {stamp(s['end'])}\n{s['text'].strip()}\n\n")
print("wrote", len(segments), "cues")

What the code does

An SRT cue is an index, a time range written as HH:MM:SS,mmm --> HH:MM:SS,mmm with a comma before the milliseconds, the text, and a blank line. That is all the function does. Set AUDIO_URL to a public HTTPS audio file of at most 10 minutes, and duration_seconds to its length; the value reserves usage, and omitting it reserves one minute. The Jobs and results page has the envelope.

The script submits with an Idempotency-Key, so a retry after a dropped connection returns the original job instead of billing a second transcription, and it polls status_url on the interval the API suggests in next_poll_after_seconds. It stops when terminal is true and then reads result_url, the same submit, poll, fetch loop the Jobs and results page describes. A 5-minute file costs $0.05 at the per-minute rate.

Cue length, pauses and offsets

Segments are gapless: each one ends exactly where the next begins, plus a default 70 ms of lead after the last word. In a subtitle file that means a cue stays up through a pause. Cap it if that bothers you, for example end = min(s['end'], s['start'] + 6). A long sentence also makes a long cue. Check the reading speed with this script and split long sentences with the word timings.

The timings come from the audio you send, not from a video. If the audio was detached from a clip, the times line up with the clip. If you chunked a long file, add each chunk's start offset before you write the file. Neither this file nor the transcript has a speaker field, because Sume STT does no diarization.

Two more details save a rejected upload. Write the file as UTF-8, which the script does; platforms refuse Latin-1 files that contain curly quotes or non-English letters. And keep the language consistent: set language_code to the language of the audio, or omit it to let the provider detect it. A wrong hint is the most common cause of a transcript full of near-misses that no timestamp function can repair. If a platform wants WebVTT instead, the change is a WEBVTT header line, a period instead of a comma in the timestamps, and no index line.

SRT or burned-in captions (read 2026-10-06)
NeedUseCost
Selectable captions on a platform that accepts SRTThis script, upload the .srt$0.01 per audio minute for the STT job
Captions burned into a vertical clipSume video captions with a style$0.20 per job for up to 60 s
BothSTT for the file, captions for the clipBoth lines above

Feeding the burned-in version

Sume's video-captions endpoint does not take an SRT upload, per the captions docs. To burn your own text, send script_text, or words, cues or segments with text, start and end. The segments from this job fit the last form, so one STT run can feed both the SRT you upload and the burned-in version, with no second transcription.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume