Sume STT to an SRT file: build subtitles from sentence segments
Sume returns timed sentence segments, not an SRT. Turn them into a valid .srt file in Python for YouTube, Vimeo or a player, with the timestamp format.

Sume speech-to-text does not output SRT or WebVTT. It returns the transcript text, word timings and, if you ask for segmentation: { mode: "sentence" }, gapless sentence segments with index, text, start and end in seconds. Writing an .srt from those segments is a loop and one timestamp function, shown below. Upload the file to YouTube, Vimeo or a player when you want selectable captions rather than burned-in ones.
Transcription is getting cheaper and more live: Microsoft's MAI-Transcribe-2-Streaming lists 60 languages at $0.54 an hour as an intro price through 2026 (read 2026-10-06 on the October tracker). Sume's batch rate is $0.01 per audio minute, or $0.60 an hour (pricing). The output is the difference: you get timings you can turn into whatever file format a platform wants.
Segments to SRT
import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]
def call(url, body=None, key=None):
h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
if key: h["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]
def stamp(t):
ms = round(t * 1000)
h, ms = divmod(ms, 3600000)
m, ms = divmod(ms, 60000)
s, ms = divmod(ms, 1000)
return f"{h:02}:{m:02}:{s:02},{ms:03}"
job = call("https://api.sume.com/v1/stt-1.0/transcribe", {
"audio_url": os.environ["AUDIO_URL"], "language_code": "en", "duration_seconds": 300,
"segmentation": {"mode": "sentence"}}, "srt-demo-v1")
while not job["terminal"]:
time.sleep(job.get("next_poll_after_seconds") or 3)
job = call(job["status_url"])
segments = call(job["result_url"])["result"]["segments"]
with open("episode.srt", "w", encoding="utf-8") as f:
for s in segments:
f.write(f"{s['index'] + 1}\n{stamp(s['start'])} --> {stamp(s['end'])}\n{s['text'].strip()}\n\n")
print("wrote", len(segments), "cues")What the code does
An SRT cue is an index, a time range written as HH:MM:SS,mmm --> HH:MM:SS,mmm with a comma before the milliseconds, the text, and a blank line. That is all the function does. Set AUDIO_URL to a public HTTPS audio file of at most 10 minutes, and duration_seconds to its length; the value reserves usage, and omitting it reserves one minute. The Jobs and results page has the envelope.
The script submits with an Idempotency-Key, so a retry after a dropped connection returns the original job instead of billing a second transcription, and it polls status_url on the interval the API suggests in next_poll_after_seconds. It stops when terminal is true and then reads result_url, the same submit, poll, fetch loop the Jobs and results page describes. A 5-minute file costs $0.05 at the per-minute rate.
Cue length, pauses and offsets
Segments are gapless: each one ends exactly where the next begins, plus a default 70 ms of lead after the last word. In a subtitle file that means a cue stays up through a pause. Cap it if that bothers you, for example end = min(s['end'], s['start'] + 6). A long sentence also makes a long cue. Check the reading speed with this script and split long sentences with the word timings.
The timings come from the audio you send, not from a video. If the audio was detached from a clip, the times line up with the clip. If you chunked a long file, add each chunk's start offset before you write the file. Neither this file nor the transcript has a speaker field, because Sume STT does no diarization.
Two more details save a rejected upload. Write the file as UTF-8, which the script does; platforms refuse Latin-1 files that contain curly quotes or non-English letters. And keep the language consistent: set language_code to the language of the audio, or omit it to let the provider detect it. A wrong hint is the most common cause of a transcript full of near-misses that no timestamp function can repair. If a platform wants WebVTT instead, the change is a WEBVTT header line, a period instead of a comma in the timestamps, and no index line.
| Need | Use | Cost |
|---|---|---|
| Selectable captions on a platform that accepts SRT | This script, upload the .srt | $0.01 per audio minute for the STT job |
| Captions burned into a vertical clip | Sume video captions with a style | $0.20 per job for up to 60 s |
| Both | STT for the file, captions for the clip | Both lines above |
Feeding the burned-in version
Sume's video-captions endpoint does not take an SRT upload, per the captions docs. To burn your own text, send script_text, or words, cues or segments with text, start and end. The segments from this job fit the last form, so one STT run can feed both the SRT you upload and the burned-in version, with no second transcription.
Sources
Related posts
More in Developers
- Preview Sume TTS sentence ids, lengths and job cost before you submit
A short Python script that splits a script like Sume's source API, groups sentences under 20,000 characters and prices each job at $0.0475 per 1,000 characters.
- Sume waitForJob pollInterval is a floor; next_poll_after_seconds wins
waitForJob never polls faster than pollInterval, and a longer next_poll_after_seconds from the server raises the gap. Defaults, timing table, sample.
- Sume webhook fails the 300-second window: find the clock drift
A valid Sume signature still fails if your server clock is more than five minutes off. Tell drift from a bad secret with a small Python check.
- Sume webhook handler over 10 seconds: acknowledge first, then work
Sume gives each webhook attempt 10 seconds. Verify, answer 2xx, then process in the background and dedupe on job_id. Node sample you can run locally.
Written by Sume