MAI-Transcribe-2-Streaming 2.5% WER: what the test audio mix is
The 2.5% word error rate comes from a chunked-streaming index: 50% AA-AgentTalk, 25% VoxPopuli, 25% Earnings22. Here is how that maps to a recorded call.

What the index is made of
The 2.5 percent word error rate quoted for MAI-Transcribe-2-Streaming is a score on one specific test mix, and that mix is built for streaming audio, not for a finished call recording. Before you carry the number into a decision about transcribing recorded calls, read what the index measures and then score your own files.
Per the Unite.AI report on Microsoft's October 1, 2026 launch (read 2026-10-07), Artificial Analysis states that its AA-WER Streaming index scores models whose audio arrives in real time, chunk by chunk, across roughly eight hours of audio drawn from three datasets. The report cites the Artificial Analysis streaming leaderboard dated September 28, 2026 for the model's 2.5 percent final word-error rate.
The snapshot gives the weights directly. Half of the score comes from a dataset the report calls AA-AgentTalk, and the rest is split evenly between VoxPopuli and Earnings22. The report describes the datasets as real-world speech with diverse accents, domain-specific language and challenging acoustic conditions.
| Dataset | Share of the index | What the report says about it |
|---|---|---|
| AA-AgentTalk | 50% | Part of the roughly eight hours of audio |
| VoxPopuli | 25% | Part of the roughly eight hours of audio |
| Earnings22 | 25% | Part of the roughly eight hours of audio |
Why a recorded call is a different test
Two details matter for a buyer. First, the audio is fed as a stream in chunks, so the score reflects a model that has to commit words before it has heard the whole sentence. A recorded file lets a transcriber use the full context. Second, the report says the latency measurements, Time to Final and Time to First Partial, start at the end of speech detected by the SileroVAD voice-activity detector. That is a timing method, and it says nothing about how fast a file transcribes.
What to take from it
Your recorded sales or support call has its own mix: your accents, your product names, your crosstalk, your phone-line compression. None of the three datasets is that mix, and the report does not claim it is. The 2.5 percent is a fair way to compare streaming models with each other under one protocol. It is not a forecast of the error rate on your audio.
Run the same question on Sume
On Sume, speech-to-text is a file job. You submit an audio URL, optionally with a language_code hint, and read back the text and a words[] array with start and end times. The public rate is $0.01 per audio minute, per the video inspect docs, which also show the same engine reached through transcribe: true on one clip. The live price is on GET /v1/catalog.
That price makes your own benchmark cheap. Pick 10 representative calls of about a minute each, transcribe them (about 10 cents at the public rate), correct a reference transcript by hand, and count the differences. The post on measuring word error rate on your own clips has a script for the counting step.
A short rule of thumb
- Take the 2.5 percent as a ranking signal among streaming models, scored on the three datasets above.
- Do not read it as a guarantee for phone audio, jargon-heavy calls or overlapping speakers.
- Keep a hand-corrected sample of your own calls as the test, and rerun it whenever you change vendor, model or settings.
- If you need live text while people talk, a streaming model is the right category. If you transcribe after the call, a file job is enough, and the streaming leaderboard is the wrong yardstick.
Sources
Related posts
More in Models
- MAI-Voice 50.3% of 4,000 listeners: how to quote the Turing claim
Microsoft said 50.3% of 4,000 listeners rated MAI-Voice as equally or more human-like than human recordings. What it covers, and a cheap test.
- Three characters, three dances: Omni IMAGE_REF and VIDEO_REF tokens
Google's Omni 1.1 demo swaps three dancers for a dog, an octopus and a bear. Here is the same request on Sume, with the 0-based reference tokens in order.
- MiniMax H3 or H3 Max after Sora: native 768p vs latent 1080p prices
On Sume, H3 renders native 480p or 768p and bills 2K/4K upscales; H3 Max adds 1080p as a latent refinement of 768p. 10-second prices side by side.
- Minimum clip length by Sume video model: 2, 3, 4 or 5 seconds
Wan 3.0 starts at 2 s, Gemini Omni Flash 1.1 at 3 s, Seedance and Genjutsu at 4 s, MiniMax H3 and Recast at 5 s. Full min and max table for a Sora port.
Written by Sume