Streaming transcripts: burn only final text into clip captions

MAI-Transcribe-2-Streaming returns partials in about 100 ms, then settled text. Burned-in captions need the final text. Filter it and pass it to Sume as cues.

6 min readSume
All posts

If you caption video with a streaming transcriber, keep only the final text and throw the partials away. Microsoft says MAI-Transcribe-2-Streaming produces its first partial hypotheses in just over 100 ms and ranks first on Artificial Analysis for both final and partial transcripts (read 2026-10-03). Partials exist so a live overlay or an agent can react mid-sentence. A word in a partial can change on the next update, and a burned-in caption cannot change once it is rendered.

This post shows the filter, the timing conversion for a clip cut from a longer stream, and the Sume request that burns the result without running speech recognition a second time.

Why partials are wrong for a rendered file

A partial is a best guess about audio that is still arriving. It is useful because it is early, and it is early because it is allowed to be wrong. A live caption overlay hides that by redrawing. A video file cannot redraw: whatever text you pass is drawn at the times you give it.

Microsoft describes the streaming model as built for low-latency, real-time transcripts in 60 languages with automatic, continuous language detection (Microsoft AI, read 2026-10-03). Its sibling model page lists captions, call analysis, accessibility and content workflows among the uses and says diarisation, timestamps and domain biasing are built in (MAI-Transcribe-2, read 2026-10-03). Neither page gives me a field-by-field event schema I can quote, so the code below uses a neutral shape of my own: each event has a text, a start and end in seconds from the stream start, and a boolean final. Map your real events onto it.

The three-step pipeline

The Sume docs describe the cues path precisely: cues or segments with text, start and end give an authored overlay that skips speech-to-text, and script_text, words, cues and segments are mutually exclusive (Video captions). Pick one.

  • Keep only events marked final, in order.
  • Shift every time by the clip's start offset and drop anything outside the clip window, trimming events that straddle an edge.
  • Send the result as cues (phrase-level) so Sume skips speech-to-text and burns exactly that copy.
Where each decision lives, read 2026-10-03.
StepDone byWhy
Partial hypotheses while speakingStreaming modelLets a live overlay or agent react early
Choosing final text and clip windowYour codeOnly settled text belongs in a file
Drawing text on video at given timesSume video captions with cuesAuthored overlay path skips speech-to-text

The filter, runnable

This builds a Sume caption request body from a list of stream events for a clip that starts at 3,600 seconds into the stream and lasts 45 seconds. The final text is the only thing that survives.

import json

events = [
    {"text": "Welcome", "start": 3600.2, "end": 3600.6, "final": False},
    {"text": "Welcome back to the show.", "start": 3600.2, "end": 3602.1, "final": True},
    {"text": "Today we", "start": 3602.4, "end": 3602.8, "final": False},
    {"text": "Today we cover pricing.", "start": 3602.4, "end": 3604.5, "final": True},
    {"text": "Next week", "start": 3650.0, "end": 3651.0, "final": True},
]
clip_start, clip_length = 3600.0, 45.0

cues = []
for e in events:
    if not e["final"]:
        continue
    start = max(e["start"] - clip_start, 0.0)
    end = min(e["end"] - clip_start, clip_length)
    if end <= 0 or start >= clip_length or end <= start:
        continue
    cues.append({"text": e["text"], "start": round(start, 2), "end": round(end, 2)})

body = {
    "video_url": "https://media.sume.com/artifacts/example/clip.mp4",
    "style": "punch",
    "cues": cues,
}
print(json.dumps(body, indent=2))

Edge cases that bite in production

Three cases come up as soon as you run this on real shows. First, an event that straddles the clip edge: the code trims it to the window, which is right for a caption but means the first words of the clip can appear mid-phrase. If that looks wrong, start the clip a beat earlier. Second, long phrases. A final event can be a whole paragraph; split it into phrases before you send it, or use the optional design.phrasing fields (max_words, max_chars, pause_seconds) to control how Sume groups words. Third, language. Continuous language detection means a bilingual host can switch mid-clip. Sume rejects Korean text on the Latin display styles slam, punch and tiktok-green with 400 caption_hangul_text_latin_style rather than drawing empty boxes, so choose a Hangul-capable style such as black-outline for any clip that may contain Hangul.

Check the burned result on a phone-size screen before you batch the rest. A caption that reads fine on a laptop can run past the safe width on a vertical clip, and a second render costs another job.

Cost and limits on the Sume side

A caption job is $0.20 for videos up to 60 seconds under the current fixed estimate, and the docs say to confirm live pricing in GET /v1/catalog. The video_url must be a fetchable public HTTPS video URL; localhost, private-network, non-HTTPS and signed or private URLs are rejected. SRT uploads are not supported, which is why the conversion above produces cues rather than a subtitle file.

One more habit saves a re-render. Keep the final events from the live run in your own storage, with the stream offsets. If you later cut a different clip from the same show, you rebuild the cues in seconds and do not pay for transcription again.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume