Voice note to SRT in Python with Sume STT sentence segments
Submit a voice note to Sume STT with sentence segmentation, poll the job, and write an SRT file from segments[]. Runnable Python with asyncio.run.

Send the recording to POST /v1/stt-1.0/transcribe with segmentation: {"mode": "sentence"}, poll until result_ready is true, then turn each entry of segments[] into an SRT cue. Segments are gapless sentences with index, text, start and end in seconds, so no word-merging logic is needed. The script below does it with only the Python standard library.
What the request and result contain
The STT request needs a public HTTPS audio_url. language_code is optional and auto-detect is the default. duration_seconds (1 to 600) improves the usage reservation; omit it and Sume reserves one minute. Word timings are always returned, so there is no flag for them.
Sentence segments are derived from those word timings. They are time ranges over your audio_url; Sume does not produce sliced files for STT.
| Field | Where | Meaning |
|---|---|---|
| audio_url | Request | Public HTTPS link to the recording |
| segmentation.mode | Request | sentence |
| duration_seconds | Request | Optional, 1 to 600, for reservation |
| text, words[] | Result | Transcript and per-word start and end |
| segments[] | Result | Gapless sentences with index, text, start, end |
The script
Set SUME_API_KEY and AUDIO_URL. The result envelope may nest the payload, so the helper unwraps data and result defensively.
import asyncio, json, os, urllib.request
API, KEY = "https://api.sume.com", os.environ["SUME_API_KEY"]
def call(method, url, body=None, idem=None):
h = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
if idem:
h["Idempotency-Key"] = idem
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(url, data=data, headers=h, method=method)
with urllib.request.urlopen(req) as r:
return json.load(r)
def stamp(sec):
ms = round(sec * 1000)
return f"{ms // 3600000:02}:{ms // 60000 % 60:02}:{ms // 1000 % 60:02},{ms % 1000:03}"
def to_srt(segs):
return "\n".join(f"{i}\n{stamp(s['start'])} --> {stamp(s['end'])}\n{s['text'].strip()}\n"
for i, s in enumerate(segs, 1))
async def main():
body = {"audio_url": os.environ["AUDIO_URL"], "segmentation": {"mode": "sentence"}, "mode": "async"}
d = await asyncio.to_thread(call, "POST", f"{API}/v1/stt-1.0/transcribe", body, "srt-note-001")
d = d.get("data", d)
while not d.get("result_ready"):
await asyncio.sleep(d.get("next_poll_after_seconds") or 2)
d = (await asyncio.to_thread(call, "GET", d["status_url"])).get("data", d)
res = await asyncio.to_thread(call, "GET", d["result_url"])
res = res.get("data", res).get("result", res)
open("note.srt", "w", encoding="utf-8").write(to_srt(res["segments"]))
asyncio.run(main())Why batch and not live
Microsoft's new MAI-Transcribe-2-Streaming sends incremental transcripts while a person speaks, and its Learn page labels it public preview without an SLA (read 2026-10-07). That is the right tool for live captions. A recorded voice note does not need partial results, and Sume's STT returns one final result per job instead.
- Jobs over 600 seconds need splitting; the schema caps
duration_secondsat 600. - Tune cue tails with
boundary_lead_ms(0 to 500, default 70). - If
segmentationcannot be built from the provider's timed words, the job fails closed with a typed error instead of returning guessed times.
Sources
Related posts
More in Developers
- Sume timeouts in one table: 30 s, 55 s, 10 s, 90 minutes
Every wait in the Sume API has its own number: sync 30 s, jobs_wait 55 s, webhook attempts 10 s, SDK helpers 10 and 20 minutes, Format runs 90 minutes.
- Sume TTS without a language field: English default, Hangul fallback
What Sume TTS 1.0 does when you leave out language, why a Hangul-only script is the one fallback, and why you should set language yourself on every job.
- Sume TTS sentence slices need wav or raw output, not mp3
Sume TTS returns per-sentence audio_url slices only for wav or raw output. With mp3 you get timings but no slices. The request that gets clips.
- Sume /v1/usage summary.final is false: a hold is open, not spent
Read GET /v1/usage?job_id= and book cost only when summary.final is true. held_usd_micros and refunded_usd_micros are not spend. Code to poll it.
Written by Sume