Live webinar captions: Sume has no streaming STT, so use chunks

MAI-Transcribe-2-Streaming is $0.54 an hour. Sume STT is $0.60, async and capped at 10 minutes a job. How to caption a webinar recording in chunks.

5 min readSume
All posts

Sume cannot caption a webinar while it is happening: its speech-to-text is a job API, not a streaming one. It takes a public audio URL, returns words with timings and optional sentence segments, and accepts at most 10 minutes of audio per job. For a live event you need a streaming recognizer; Microsoft's MAI-Transcribe-2-Streaming lists 60 languages at $0.54 an hour as an introductory price through the end of 2026 (read 2026-10-06 on the October tracker). For the replay, the clip library and the social cut-downs, Sume does the job at $0.01 per audio minute, which is $0.60 an hour (pricing, read 2026-10-06).

That makes the honest answer a split. Use a streaming tool for the live caption window. Then run the recording through Sume for a clean transcript with timings, searchable chapters and burned-in captions on the highlight clips.

Chunk, transcribe, offset

Cut the recording into pieces of 10 minutes or less, upload them where Sume can fetch them over public HTTPS, and submit one STT job per piece with duration_seconds set to its length. The reservation is made from that field; omit it and Sume reserves one minute. Run the jobs in parallel, then add each piece's start offset to its segment times. The script below does it for any list of piece URLs and prints a time-coded transcript.

import json, os, time, urllib.request as u
from concurrent.futures import ThreadPoolExecutor

KEY = os.environ["SUME_API_KEY"]
CHUNKS = [("https://example.com/webinar-part-1.mp3", 600), ("https://example.com/webinar-part-2.mp3", 540)]

def call(url, body=None, key=None):
    h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
    if key: h["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body else None
    return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]

def transcribe(i, url, seconds):
    job = call("https://api.sume.com/v1/stt-1.0/transcribe",
               {"audio_url": url, "language_code": "en", "duration_seconds": seconds,
                "segmentation": {"mode": "sentence"}}, f"webinar-v1-{i}")
    while not job["terminal"]:
        time.sleep(job.get("next_poll_after_seconds") or 3)
        job = call(job["status_url"])
    return call(job["result_url"])["result"]["segments"]

with ThreadPoolExecutor(4) as pool:
    parts = list(pool.map(lambda c: transcribe(c[0], *c[1]), enumerate(CHUNKS)))
offset = 0
for (url, seconds), segs in zip(CHUNKS, parts):
    for s in segs:
        t = offset + s["start"]
        print(f"[{int(t // 60):02d}:{int(t % 60):02d}] {s['text']}")
    offset += seconds

What comes back

Replace the two example URLs with your own pieces; ffmpeg -f segment -segment_time 600 makes ten-minute files in one command. The job result includes segments only when you pass segmentation.mode: "sentence", and each segment has index, text, start and end. Word timings come back on every job. There is no diarization and no speaker field, so a panel needs speaker names added by hand. Documented shapes are in Jobs and results.

90-minute webinar, two routes (read 2026-10-06)
RouteLive captionsReplay transcriptCost for 90 minutes
MAI-Transcribe-2-Streaming (tracker)Yes, streamingYes$0.81 at the $0.54 intro price
Sume STT, 9 chunks of 10 minutesNoYes, with timings and sentences$0.90 at $0.01 per minute
BothYesYes, from Sume$1.71

From the transcript to captioned clips

From the transcript to video: cut each highlight to 60 seconds or less and run it through video captions. That job is $0.20 per clip of up to 60 seconds, takes a style, and can use your own script_text so product names are spelled the way you want them. The words from the chunk run give you the cut points; the caption job transcribes the clip again on its own.

Two practical notes on the chunks. Cut on silence if you can, so no word is split across two files; a word cut in half is transcribed badly on both sides, and the offset arithmetic in the script cannot repair it. And keep the duration_seconds value honest. It drives the reservation, so a 600 on a 90-second clip holds more balance than the job needs, while omitting it reserves one minute and fits a short clip only. Neither changes what the finished job transcribes, but both change what your balance shows while the batch runs.

If the recording has an interpreter or a second language, set language_code per chunk instead of one value for the whole event. Omit it for auto-detect on a chunk where you do not know. A language hint belongs to a single job, so a bilingual panel is best cut at the language switch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume