MAI-Transcribe-2-Streaming 0.13 s to final: when the clock starts
The 0.13 second figure is measured from end of speech found by a VAD, and partials arrive in about 100 ms. Why neither number is the wait for a file transcript.

Short answer
The 0.13 seconds you will see attached to MAI-Transcribe-2-Streaming is the time from the end of a spoken utterance to the final transcript, and the clock starts when a voice-activity detector says speech has stopped. It is not the time a file takes to transcribe, and it cannot be compared with a batch job's wait.
All figures here come from the Unite.AI report on Microsoft's October 1, 2026 launch (read 2026-10-07).
The report describes three different time-related claims, and they are easy to blur together.
| Statement | Source in the report | Clock starts at |
|---|---|---|
| First hypotheses (partials) in just over 100 ms of receiving audio | Microsoft | Audio arriving at the model |
| 0.13 s to final transcription | Artificial Analysis leaderboard chart dated Sept 28, 2026 | End of speech detected by SileroVAD |
| Words appear twice as fast as the closest competitor | Microsoft internal evaluations | Not specified in the report |
Why the start line matters
The second row is the one most often quoted, and its start line is the subtle part. Artificial Analysis, as quoted by the report, starts both Time to Final and Time to First Partial at the end of speech found by the SileroVAD detector. A system that waits longer to decide you have stopped talking is not penalized by that number, because the detector, not the system under test, sets the starting line.
That is a reasonable way to compare models on equal terms. It also means the number describes the model's tail latency after an utterance, inside a live conversation. It does not include your network, your own end-of-turn logic or the time to speak a reply.
What to measure for a recorded file
A recorded file has no live conversation around it, so the useful question is different: how long from submitting this file to having a transcript I can use. That depends on file length, queue time and your polling, none of which the streaming figures measure.
Sume's speech-to-text is built around that file question. You submit an audio URL, and read the transcript and words[] timings when the job is ready. The docs for video inspect, which can run the same transcription on a clip, state a 30-second synchronous wait, after which you get a queued job and poll instead. Treat those as the knobs to measure: time your own files from submit to result, as the post on benchmarking latency yourself shows for speech output.
Which number to use when
- Building a voice agent that replies mid-conversation: use the streaming figures, and test end-of-turn handling yourself.
- Transcribing recorded calls, interviews or clips: ignore the 0.13 s and compare accuracy and price per audio minute.
- Burning captions after the fact: the latency does not matter, only that the word timings line up.
Questions to ask any vendor's latency number
Whatever the vendor, ask four things before comparing latency figures. Where does the clock start? What does it stop on, a partial or a final transcript? Was the audio sent as a stream or as a file? Whose network was in the path?
The launch report answers the first two for this model: the clock starts at end of speech from a voice-activity detector, and there are two stop points, first partial and final. Treat the third and fourth as open unless the vendor says.
- Where the clock starts.
- What it stops on.
- Stream or file.
- Which network was in the path.
Sources
Related posts
More in Models
- MAI-Transcribe-2-Streaming 2.5% WER: what the test audio mix is
The 2.5% word error rate comes from a chunked-streaming index: 50% AA-AgentTalk, 25% VoxPopuli, 25% Earnings22. Here is how that maps to a recorded call.
- MAI-Voice 50.3% of 4,000 listeners: how to quote the Turing claim
Microsoft said 50.3% of 4,000 listeners rated MAI-Voice as equally or more human-like than human recordings. What it covers, and a cheap test.
- Three characters, three dances: Omni IMAGE_REF and VIDEO_REF tokens
Google's Omni 1.1 demo swaps three dancers for a dog, an octopus and a bear. Here is the same request on Sume, with the 0-based reference tokens in order.
- MiniMax H3 or H3 Max after Sora: native 768p vs latent 1080p prices
On Sume, H3 renders native 480p or 768p and bills 2K/4K upscales; H3 Max adds 1080p as a latent refinement of 768p. 10-second prices side by side.
Written by Sume