Streaming transcripts: burn only final text into clip captions
MAI-Transcribe-2-Streaming returns partials in about 100 ms, then settled text. Burned-in captions need the final text. Filter it and pass it to Sume as cues.

If you caption video with a streaming transcriber, keep only the final text and throw the partials away. Microsoft says MAI-Transcribe-2-Streaming produces its first partial hypotheses in just over 100 ms and ranks first on Artificial Analysis for both final and partial transcripts (read 2026-10-03). Partials exist so a live overlay or an agent can react mid-sentence. A word in a partial can change on the next update, and a burned-in caption cannot change once it is rendered.
This post shows the filter, the timing conversion for a clip cut from a longer stream, and the Sume request that burns the result without running speech recognition a second time.
Why partials are wrong for a rendered file
A partial is a best guess about audio that is still arriving. It is useful because it is early, and it is early because it is allowed to be wrong. A live caption overlay hides that by redrawing. A video file cannot redraw: whatever text you pass is drawn at the times you give it.
Microsoft describes the streaming model as built for low-latency, real-time transcripts in 60 languages with automatic, continuous language detection (Microsoft AI, read 2026-10-03). Its sibling model page lists captions, call analysis, accessibility and content workflows among the uses and says diarisation, timestamps and domain biasing are built in (MAI-Transcribe-2, read 2026-10-03). Neither page gives me a field-by-field event schema I can quote, so the code below uses a neutral shape of my own: each event has a text, a start and end in seconds from the stream start, and a boolean final. Map your real events onto it.
The three-step pipeline
The Sume docs describe the cues path precisely: cues or segments with text, start and end give an authored overlay that skips speech-to-text, and script_text, words, cues and segments are mutually exclusive (Video captions). Pick one.
- Keep only events marked final, in order.
- Shift every time by the clip's start offset and drop anything outside the clip window, trimming events that straddle an edge.
- Send the result as
cues(phrase-level) so Sume skips speech-to-text and burns exactly that copy.
| Step | Done by | Why |
|---|---|---|
| Partial hypotheses while speaking | Streaming model | Lets a live overlay or agent react early |
| Choosing final text and clip window | Your code | Only settled text belongs in a file |
| Drawing text on video at given times | Sume video captions with cues | Authored overlay path skips speech-to-text |
The filter, runnable
This builds a Sume caption request body from a list of stream events for a clip that starts at 3,600 seconds into the stream and lasts 45 seconds. The final text is the only thing that survives.
import json
events = [
{"text": "Welcome", "start": 3600.2, "end": 3600.6, "final": False},
{"text": "Welcome back to the show.", "start": 3600.2, "end": 3602.1, "final": True},
{"text": "Today we", "start": 3602.4, "end": 3602.8, "final": False},
{"text": "Today we cover pricing.", "start": 3602.4, "end": 3604.5, "final": True},
{"text": "Next week", "start": 3650.0, "end": 3651.0, "final": True},
]
clip_start, clip_length = 3600.0, 45.0
cues = []
for e in events:
if not e["final"]:
continue
start = max(e["start"] - clip_start, 0.0)
end = min(e["end"] - clip_start, clip_length)
if end <= 0 or start >= clip_length or end <= start:
continue
cues.append({"text": e["text"], "start": round(start, 2), "end": round(end, 2)})
body = {
"video_url": "https://media.sume.com/artifacts/example/clip.mp4",
"style": "punch",
"cues": cues,
}
print(json.dumps(body, indent=2))
Edge cases that bite in production
Three cases come up as soon as you run this on real shows. First, an event that straddles the clip edge: the code trims it to the window, which is right for a caption but means the first words of the clip can appear mid-phrase. If that looks wrong, start the clip a beat earlier. Second, long phrases. A final event can be a whole paragraph; split it into phrases before you send it, or use the optional design.phrasing fields (max_words, max_chars, pause_seconds) to control how Sume groups words. Third, language. Continuous language detection means a bilingual host can switch mid-clip. Sume rejects Korean text on the Latin display styles slam, punch and tiktok-green with 400 caption_hangul_text_latin_style rather than drawing empty boxes, so choose a Hangul-capable style such as black-outline for any clip that may contain Hangul.
Check the burned result on a phone-size screen before you batch the rest. A caption that reads fine on a laptop can run past the safe width on a vertical clip, and a second render costs another job.
Cost and limits on the Sume side
A caption job is $0.20 for videos up to 60 seconds under the current fixed estimate, and the docs say to confirm live pricing in GET /v1/catalog. The video_url must be a fetchable public HTTPS video URL; localhost, private-network, non-HTTPS and signed or private URLs are rejected. SRT uploads are not supported, which is why the conversion above produces cues rather than a subtitle file.
One more habit saves a re-render. Keep the final events from the live run in your own storage, with the stream offsets. If you later cut a different clip from the same show, you rebuild the cues in seconds and do not pay for transcription again.
Sources
Related posts
More in Developers
- STT sentence segmentation: unpunctuated speech split on silence
Sume STT 1.0 segmentation mode sentence returns gapless sentence segments, splitting unpunctuated runs on silence. Options, boundary_lead_ms and caveats.
- Submit 20 music takes at once: queued is normal, queue_full is not
Sume accepts music jobs as queued while the queue has room, then returns 429 queue_full. Default accepted-job capacity by plan, from 6 on Free to 120 on Scale.
- subscribeFormatRun onCreated: save the run id, set your own key
subscribeFormatRun makes a new Idempotency-Key per call unless you pass one. Save the run id in onCreated and pass a stable key so a restart cannot double-run.
- Sume API rate limits by plan: requests per minute for writes and reads
Sume gives every API key a per-minute budget set by plan: 120 writes on Free up to 1200 on Scale, with reads at forty times the write number. Table and headers.
Written by Sume