MAI-Transcribe-2-Streaming 0.13 s to final: when the clock starts

The 0.13 second figure is measured from end of speech found by a VAD, and partials arrive in about 100 ms. Why neither number is the wait for a file transcript.

5 min readSume
All posts

Short answer

The 0.13 seconds you will see attached to MAI-Transcribe-2-Streaming is the time from the end of a spoken utterance to the final transcript, and the clock starts when a voice-activity detector says speech has stopped. It is not the time a file takes to transcribe, and it cannot be compared with a batch job's wait.

All figures here come from the Unite.AI report on Microsoft's October 1, 2026 launch (read 2026-10-07).

The report describes three different time-related claims, and they are easy to blur together.

Latency statements in the launch report (read 2026-10-07)
StatementSource in the reportClock starts at
First hypotheses (partials) in just over 100 ms of receiving audioMicrosoftAudio arriving at the model
0.13 s to final transcriptionArtificial Analysis leaderboard chart dated Sept 28, 2026End of speech detected by SileroVAD
Words appear twice as fast as the closest competitorMicrosoft internal evaluationsNot specified in the report

Why the start line matters

The second row is the one most often quoted, and its start line is the subtle part. Artificial Analysis, as quoted by the report, starts both Time to Final and Time to First Partial at the end of speech found by the SileroVAD detector. A system that waits longer to decide you have stopped talking is not penalized by that number, because the detector, not the system under test, sets the starting line.

That is a reasonable way to compare models on equal terms. It also means the number describes the model's tail latency after an utterance, inside a live conversation. It does not include your network, your own end-of-turn logic or the time to speak a reply.

What to measure for a recorded file

A recorded file has no live conversation around it, so the useful question is different: how long from submitting this file to having a transcript I can use. That depends on file length, queue time and your polling, none of which the streaming figures measure.

Sume's speech-to-text is built around that file question. You submit an audio URL, and read the transcript and words[] timings when the job is ready. The docs for video inspect, which can run the same transcription on a clip, state a 30-second synchronous wait, after which you get a queued job and poll instead. Treat those as the knobs to measure: time your own files from submit to result, as the post on benchmarking latency yourself shows for speech output.

Which number to use when

  • Building a voice agent that replies mid-conversation: use the streaming figures, and test end-of-turn handling yourself.
  • Transcribing recorded calls, interviews or clips: ignore the 0.13 s and compare accuracy and price per audio minute.
  • Burning captions after the fact: the latency does not matter, only that the word timings line up.

Questions to ask any vendor's latency number

Whatever the vendor, ask four things before comparing latency figures. Where does the clock start? What does it stop on, a partial or a final transcript? Was the audio sent as a stream or as a file? Whose network was in the path?

The launch report answers the first two for this model: the clock starts at end of speech from a voice-activity detector, and there are two stop points, first partial and final. Treat the third and fourth as open unless the vendor says.

  • Where the clock starts.
  • What it stops on.
  • Stream or file.
  • Which network was in the path.

Sources

Related posts

More in Models

All Models posts

Written by Sume