MAI-Transcribe-2-Streaming '2x faster': measure your caption delay
Microsoft says words appear 2x faster than its closest competitor. Measure your own delay from speech to caption with a p50 and p95 script before you switch.

Microsoft says words appear in the transcript 2x faster with MAI-Transcribe-2-Streaming than with its closest competitor, with first partial hypotheses in just over 100 ms (read 2026-10-03). Treat that as a reason to test, not a number to plan around. The delay your viewers feel is the gap between a word being spoken and being on screen, measured on your audio, your network and your overlay, at the tail, not the average.
This post defines three delays, gives a script that turns logged timestamps into p50 and p95, and shows where a finished-file caption job, such as Sume's, sits relative to a live one.
Three delays, not one
The marketing number is usually the first row. A viewer reads the third. Between them sit network, your own buffering, any smoothing you add to stop words flickering, and the render loop.
| Delay | From | To | Vendor claim |
|---|---|---|---|
| First partial | Audio of a word sent | First partial containing it appears | Just over 100 ms (Microsoft) |
| Settled text | Audio of a word sent | Final text for that phrase | Not stated as a number on the pages I read |
| On-screen | Audio of a word sent | Pixels drawn in your overlay | Yours to measure |
What a claim of '2x' needs from you
A relative claim is only as good as its baseline, which is 'our closest competitor' on Microsoft's tests. To check it for you, run the same recorded audio through the streaming model and the engine you use now, log timestamps on your side at the moment each chunk of audio leaves and each result arrives, and compare both engines on the same clock. Use at least a few hours of real speech in your real languages, because latency and accuracy both move with accents and noise. Microsoft also says the model ranks first on Artificial Analysis for accuracy on final and partial transcripts; accuracy and delay trade off, so record word error alongside delay, or a faster engine that drops words will look better than it is.
From timestamps to p50 and p95
Log two timestamps per word or phrase: when its audio was sent and when its text arrived. This script computes the delay distribution and flags the worst cases. The sample data stands in for your logs.
import statistics
# (audio_sent_seconds, text_arrived_seconds) per phrase, from your own logs
samples = [
(10.00, 10.14), (11.20, 11.31), (12.00, 12.46), (13.50, 13.62),
(14.10, 14.25), (15.00, 15.90), (16.30, 16.41), (17.00, 17.13),
(18.20, 18.33), (19.00, 19.20),
]
delays = sorted(round(arrived - sent, 3) for sent, arrived in samples)
p50 = statistics.median(delays)
p95 = delays[int(0.95 * (len(delays) - 1))]
print("n=%d p50=%.0f ms p95=%.0f ms max=%.0f ms" % (len(delays), p50 * 1000, p95 * 1000, delays[-1] * 1000))
slow = [d for d in delays if d > 0.5]
print("phrases over 500 ms:", len(slow))
A test plan you can finish in an afternoon
Pick six recordings: two clean studio reads, two noisy rooms, one with two speakers overlapping, one that switches language. Replay each into every engine at real-time speed, since a streaming model fed faster than real time tells you nothing about live delay. Log the send and arrival times on one machine so the clocks agree. Score the final text against a human transcript for word errors.
Then write down the decision before you look at the result: for example, switch only if p95 on-screen delay improves by at least 30 percent and word error does not rise. A rule fixed in advance stops a good-looking chart from deciding for you. If the new engine passes, keep the recordings and the script, because the vendor will ship another model next quarter and you will want to run the same test again.
Where a finished-file caption job fits
Not everything needs to be live. When the video is already finished, Sume's video captions is a job: you give it a public HTTPS video URL, it returns a captioned video, and you poll the job through Jobs and results. The delay that matters there is submit to completed, which you can time with the same method. It is a poor fit for a live stream and a good fit for the clip you post afterwards, because a fixed $0.20 per job for videos up to 60 seconds is easy to budget.
A practical decision rule: if the viewer is watching while the speaker talks, you need a streaming transcriber and you should measure its p95. If the viewer watches a file later, measure how long your caption step takes and what it costs, and spend your attention on the style and the wording instead.
Sources
Related posts
More in Developers
- Microsoft Agent Framework hosted MCP tool for Sume (Python)
Attach https://mcp.sume.com/mcp to an Agent Framework agent with get_mcp_tool, allowed_tools and approval_mode. Foundry runs the calls, so mind the key.
- Migrate Video 1.0 to sume/auto on POST /v1/videos, field by field
Video 1.0 is retiring soon. Which fields move, which are ignored or rejected, and the request to send on POST /v1/videos with sume/auto.
- Migrate POST /v1/video-router/generate to /v1/videos: a field map
Same model ids, same jobs, different wire. How image_url and reference_image_urls become frame_images and input_references, and what stays on Video Router.
- MiniMax H3 Max in TypeScript: submit, poll and download with fetch
A TypeScript script under 30 lines: submit a minimax-h3-max job on Sume, poll until completed, save the MP4. Status values and the 409 on failed jobs.
Written by Sume