Voice note to SRT in Python with Sume STT sentence segments

Submit a voice note to Sume STT with sentence segmentation, poll the job, and write an SRT file from segments[]. Runnable Python with asyncio.run.

5 min readSume
All posts

Send the recording to POST /v1/stt-1.0/transcribe with segmentation: {"mode": "sentence"}, poll until result_ready is true, then turn each entry of segments[] into an SRT cue. Segments are gapless sentences with index, text, start and end in seconds, so no word-merging logic is needed. The script below does it with only the Python standard library.

What the request and result contain

The STT request needs a public HTTPS audio_url. language_code is optional and auto-detect is the default. duration_seconds (1 to 600) improves the usage reservation; omit it and Sume reserves one minute. Word timings are always returned, so there is no flag for them.

Sentence segments are derived from those word timings. They are time ranges over your audio_url; Sume does not produce sliced files for STT.

STT 1.0 request and result fields used here - Sume OpenAPI (read 2026-10-07)
FieldWhereMeaning
audio_urlRequestPublic HTTPS link to the recording
segmentation.modeRequestsentence
duration_secondsRequestOptional, 1 to 600, for reservation
text, words[]ResultTranscript and per-word start and end
segments[]ResultGapless sentences with index, text, start, end

The script

Set SUME_API_KEY and AUDIO_URL. The result envelope may nest the payload, so the helper unwraps data and result defensively.

import asyncio, json, os, urllib.request
API, KEY = "https://api.sume.com", os.environ["SUME_API_KEY"]
def call(method, url, body=None, idem=None):
    h = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
    if idem:
        h["Idempotency-Key"] = idem
    data = json.dumps(body).encode() if body is not None else None
    req = urllib.request.Request(url, data=data, headers=h, method=method)
    with urllib.request.urlopen(req) as r:
        return json.load(r)
def stamp(sec):
    ms = round(sec * 1000)
    return f"{ms // 3600000:02}:{ms // 60000 % 60:02}:{ms // 1000 % 60:02},{ms % 1000:03}"
def to_srt(segs):
    return "\n".join(f"{i}\n{stamp(s['start'])} --> {stamp(s['end'])}\n{s['text'].strip()}\n"
                     for i, s in enumerate(segs, 1))
async def main():
    body = {"audio_url": os.environ["AUDIO_URL"], "segmentation": {"mode": "sentence"}, "mode": "async"}
    d = await asyncio.to_thread(call, "POST", f"{API}/v1/stt-1.0/transcribe", body, "srt-note-001")
    d = d.get("data", d)
    while not d.get("result_ready"):
        await asyncio.sleep(d.get("next_poll_after_seconds") or 2)
        d = (await asyncio.to_thread(call, "GET", d["status_url"])).get("data", d)
    res = await asyncio.to_thread(call, "GET", d["result_url"])
    res = res.get("data", res).get("result", res)
    open("note.srt", "w", encoding="utf-8").write(to_srt(res["segments"]))
asyncio.run(main())

Why batch and not live

Microsoft's new MAI-Transcribe-2-Streaming sends incremental transcripts while a person speaks, and its Learn page labels it public preview without an SLA (read 2026-10-07). That is the right tool for live captions. A recorded voice note does not need partial results, and Sume's STT returns one final result per job instead.

  • Jobs over 600 seconds need splitting; the schema caps duration_seconds at 600.
  • Tune cue tails with boundary_lead_ms (0 to 500, default 70).
  • If segmentation cannot be built from the provider's timed words, the job fails closed with a typed error instead of returning guessed times.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume