Test MAI-Transcribe-2-Streaming's accuracy claim on your own audio

Microsoft says its streaming transcriber ranks first on Artificial Analysis. A 10-clip test on your own audio, with a word error rate script and Sume STT.

6 min readSume
All posts

Should you believe a first place on a public speech-to-text leaderboard for your own audio? Test it. A leaderboard averages other people's recordings; your calls, accents and product names may behave differently. This post gives a small test you can run on any transcriber, with a word error rate script that needs no API.

Microsoft's announcement (read 2026-10-04) says MAI-Transcribe-2-Streaming ranks first on Artificial Analysis for accuracy on both final and partial transcripts, produces first text just over 100 ms after receiving audio, covers 60 languages with continuous language detection, and costs $0.54 per hour of audio through year-end. Those are Microsoft's statements; none of them is a measurement of your audio.

Build the test set

Pick ten clips of 30 to 60 seconds that look like production: noisy ones, fast talkers, jargon, and at least one with a name your brand cares about. Type a reference transcript for each by hand. Do not start from a machine transcript, or you will grade the machine against itself.

  • Keep punctuation and casing out of the comparison; compare lowercase words only.
  • Write numbers the way you want them to appear, then decide up front whether twenty and 20 count as a match.
  • Hold back two clips you never tune on.

Where Sume fits

Sume cannot run MAI-Transcribe-2-Streaming; its STT surface uses a provider the public contract keeps internal. What Sume can do is prepare audio and give you one more candidate to score. For a video source, audio detach returns a wav with channels: "mono" and sample_rate: 16000, the STT shape, at $0.01 per job. Then POST /v1/stt-1.0/transcribe returns text and words[] with start and end seconds.

Video inspect documents the STT rate as $0.01 per audio minute and says language_code is a hint you can omit for auto-detect. Confirm live prices in GET /v1/catalog before you budget.

Score it

Word error rate is substitutions plus deletions plus insertions, divided by the number of reference words. This script computes it for one pair of strings; run it per clip for each candidate and compare the totals.

import re

def words(s):
    return re.findall(r"[a-z0-9']+", s.lower())

def wer(ref, hyp):
    r, h = words(ref), words(hyp)
    d = list(range(len(h) + 1))
    for i in range(1, len(r) + 1):
        prev, d[0] = d[0], i
        for j in range(1, len(h) + 1):
            cur = min(d[j] + 1, d[j - 1] + 1,
                      prev + (r[i - 1] != h[j - 1]))
            prev, d[j] = d[j], cur
    return d[len(h)] / max(len(r), 1)

print(round(wer("book a table for two at eight", "book a table for you at eight"), 3))

Read the result honestly

Ten short clips is a smoke test, not a benchmark. If two candidates land within a point or two of each other, the difference is noise, and price, language coverage and latency should decide. If one is far worse on your hard clips, that is the finding that matters, whatever the leaderboard says.

Streaming and file jobs also differ in shape: a streaming model returns partial text while audio arrives, and Sume STT is a job you poll. See live versus file transcription before you compare their prices.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume