MAI-Transcribe-2-Streaming '2x faster': measure your caption delay

Microsoft says words appear 2x faster than its closest competitor. Measure your own delay from speech to caption with a p50 and p95 script before you switch.

6 min readSume
All posts

Microsoft says words appear in the transcript 2x faster with MAI-Transcribe-2-Streaming than with its closest competitor, with first partial hypotheses in just over 100 ms (read 2026-10-03). Treat that as a reason to test, not a number to plan around. The delay your viewers feel is the gap between a word being spoken and being on screen, measured on your audio, your network and your overlay, at the tail, not the average.

This post defines three delays, gives a script that turns logged timestamps into p50 and p95, and shows where a finished-file caption job, such as Sume's, sits relative to a live one.

Three delays, not one

The marketing number is usually the first row. A viewer reads the third. Between them sit network, your own buffering, any smoothing you add to stop words flickering, and the render loop.

Definitions for a measurement plan; vendor claims from Microsoft AI's launch post, read 2026-10-03.
DelayFromToVendor claim
First partialAudio of a word sentFirst partial containing it appearsJust over 100 ms (Microsoft)
Settled textAudio of a word sentFinal text for that phraseNot stated as a number on the pages I read
On-screenAudio of a word sentPixels drawn in your overlayYours to measure

What a claim of '2x' needs from you

A relative claim is only as good as its baseline, which is 'our closest competitor' on Microsoft's tests. To check it for you, run the same recorded audio through the streaming model and the engine you use now, log timestamps on your side at the moment each chunk of audio leaves and each result arrives, and compare both engines on the same clock. Use at least a few hours of real speech in your real languages, because latency and accuracy both move with accents and noise. Microsoft also says the model ranks first on Artificial Analysis for accuracy on final and partial transcripts; accuracy and delay trade off, so record word error alongside delay, or a faster engine that drops words will look better than it is.

From timestamps to p50 and p95

Log two timestamps per word or phrase: when its audio was sent and when its text arrived. This script computes the delay distribution and flags the worst cases. The sample data stands in for your logs.

import statistics

# (audio_sent_seconds, text_arrived_seconds) per phrase, from your own logs
samples = [
    (10.00, 10.14), (11.20, 11.31), (12.00, 12.46), (13.50, 13.62),
    (14.10, 14.25), (15.00, 15.90), (16.30, 16.41), (17.00, 17.13),
    (18.20, 18.33), (19.00, 19.20),
]

delays = sorted(round(arrived - sent, 3) for sent, arrived in samples)
p50 = statistics.median(delays)
p95 = delays[int(0.95 * (len(delays) - 1))]
print("n=%d p50=%.0f ms p95=%.0f ms max=%.0f ms" % (len(delays), p50 * 1000, p95 * 1000, delays[-1] * 1000))
slow = [d for d in delays if d > 0.5]
print("phrases over 500 ms:", len(slow))

A test plan you can finish in an afternoon

Pick six recordings: two clean studio reads, two noisy rooms, one with two speakers overlapping, one that switches language. Replay each into every engine at real-time speed, since a streaming model fed faster than real time tells you nothing about live delay. Log the send and arrival times on one machine so the clocks agree. Score the final text against a human transcript for word errors.

Then write down the decision before you look at the result: for example, switch only if p95 on-screen delay improves by at least 30 percent and word error does not rise. A rule fixed in advance stops a good-looking chart from deciding for you. If the new engine passes, keep the recordings and the script, because the vendor will ship another model next quarter and you will want to run the same test again.

Where a finished-file caption job fits

Not everything needs to be live. When the video is already finished, Sume's video captions is a job: you give it a public HTTPS video URL, it returns a captioned video, and you poll the job through Jobs and results. The delay that matters there is submit to completed, which you can time with the same method. It is a poor fit for a live stream and a good fit for the clip you post afterwards, because a fixed $0.20 per job for videos up to 60 seconds is easy to budget.

A practical decision rule: if the viewer is watching while the speaker talks, you need a streaming transcriber and you should measure its p95. If the viewer watches a file later, measure how long your caption step takes and what it costs, and spend your attention on the style and the wording instead.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume