MAI-Transcribe-2-Streaming 2.5% WER: what the test audio mix is

The 2.5% word error rate comes from a chunked-streaming index: 50% AA-AgentTalk, 25% VoxPopuli, 25% Earnings22. Here is how that maps to a recorded call.

5 min readSume
All posts

What the index is made of

The 2.5 percent word error rate quoted for MAI-Transcribe-2-Streaming is a score on one specific test mix, and that mix is built for streaming audio, not for a finished call recording. Before you carry the number into a decision about transcribing recorded calls, read what the index measures and then score your own files.

Per the Unite.AI report on Microsoft's October 1, 2026 launch (read 2026-10-07), Artificial Analysis states that its AA-WER Streaming index scores models whose audio arrives in real time, chunk by chunk, across roughly eight hours of audio drawn from three datasets. The report cites the Artificial Analysis streaming leaderboard dated September 28, 2026 for the model's 2.5 percent final word-error rate.

The snapshot gives the weights directly. Half of the score comes from a dataset the report calls AA-AgentTalk, and the rest is split evenly between VoxPopuli and Earnings22. The report describes the datasets as real-world speech with diverse accents, domain-specific language and challenging acoustic conditions.

AA-WER Streaming index composition as reported by Unite.AI (read 2026-10-07)
DatasetShare of the indexWhat the report says about it
AA-AgentTalk50%Part of the roughly eight hours of audio
VoxPopuli25%Part of the roughly eight hours of audio
Earnings2225%Part of the roughly eight hours of audio

Why a recorded call is a different test

Two details matter for a buyer. First, the audio is fed as a stream in chunks, so the score reflects a model that has to commit words before it has heard the whole sentence. A recorded file lets a transcriber use the full context. Second, the report says the latency measurements, Time to Final and Time to First Partial, start at the end of speech detected by the SileroVAD voice-activity detector. That is a timing method, and it says nothing about how fast a file transcribes.

What to take from it

Your recorded sales or support call has its own mix: your accents, your product names, your crosstalk, your phone-line compression. None of the three datasets is that mix, and the report does not claim it is. The 2.5 percent is a fair way to compare streaming models with each other under one protocol. It is not a forecast of the error rate on your audio.

Run the same question on Sume

On Sume, speech-to-text is a file job. You submit an audio URL, optionally with a language_code hint, and read back the text and a words[] array with start and end times. The public rate is $0.01 per audio minute, per the video inspect docs, which also show the same engine reached through transcribe: true on one clip. The live price is on GET /v1/catalog.

That price makes your own benchmark cheap. Pick 10 representative calls of about a minute each, transcribe them (about 10 cents at the public rate), correct a reference transcript by hand, and count the differences. The post on measuring word error rate on your own clips has a script for the counting step.

A short rule of thumb

  • Take the 2.5 percent as a ranking signal among streaming models, scored on the three datasets above.
  • Do not read it as a guarantee for phone audio, jargon-heavy calls or overlapping speakers.
  • Keep a hand-corrected sample of your own calls as the test, and rerun it whenever you change vendor, model or settings.
  • If you need live text while people talk, a streaming model is the right category. If you transcribe after the call, a file job is enough, and the streaming leaderboard is the wrong yardstick.

Sources

Related posts

More in Models

All Models posts

Written by Sume