Score your own audio: a word error rate script for Sume STT

Microsoft quotes 5.2% WER on FLEURS for MAI-Transcribe-2. Measure your audio: a short Python WER function, a reference file and a cent-a-minute Sume STT run.

6 min readSume
All posts

Microsoft's MAI-Transcribe-2 post states an "average Word-Error-Rate of 5.2%" on the FLEURS benchmark and ranks it second on the Artificial Analysis word-error-rate leaderboard; the streaming post ranks MAI-Transcribe-2-Streaming first on that site for final and partial transcripts (read 2026-10-08). Those are public-benchmark numbers. To know what a model does on your calls, type a reference transcript for 10 to 20 clips and compute word error rate yourself.

Below is a dependency-free Python function for it, followed by a run plan using Sume STT at about a cent per audio minute.

How to read the claims

FLEURS is read speech in many languages; a call-center line with crosstalk is a different task. A rank is relative to the other models on one site's test mix on a date. The pages I read give no per-language or per-domain breakdown, so a 5.2% average does not tell you your own rate. Treat vendor figures as a reason to test, not as the test.

What the Microsoft pages claim, and what each does not tell you (read 2026-10-08)
ClaimPageNot stated
5.2% average WER on FLEURSMAI-Transcribe-2 postPer-language and noisy-audio results
#2 on Artificial Analysis WERMAI-Transcribe-2 postDate and test mix
#1 final and partial, streamingStreaming postWER value
60 languagesBoth postsAccuracy per language

The function

Word error rate is substitutions plus deletions plus insertions, divided by the number of reference words. Normalize case and punctuation first, and decide how you treat numbers ("100" versus "one hundred") before you start: both vendors' output styles differ, and an unnormalized comparison punishes the wrong thing.

import re

def words(s):
    return re.sub(r"[^\w\s']", " ", s.lower()).split()

def wer(ref, hyp):
    r, h = words(ref), words(hyp)
    d = list(range(len(h) + 1))
    for i in range(1, len(r) + 1):
        prev, d[0] = d[0], i
        for j in range(1, len(h) + 1):
            cur = d[j]
            d[j] = min(d[j] + 1, d[j - 1] + 1,
                       prev + (r[i - 1] != h[j - 1]))
            prev = cur
    return d[len(h)] / max(1, len(r))

print(wer("we ship on Friday", "we shipped on friday"))

Run plan on Sume

Pick 15 clips of 60 seconds from your real traffic, transcribe each by hand once, and send each to POST /v1/stt-1.0/transcribe with duration_seconds: 60. That is 15 jobs at one cent, 15 cents. Read text from the job result, run wer, and average. Do the same through any other vendor's API, with the same normalization.

Sume's STT result also carries word timings, so you can see where errors cluster in time. It returns no speaker labels, so score each speaker's channel as a separate file if the recording has two sides.

Reading the result

Report WER per clip and as a pooled figure (total errors over total reference words), because a short noisy clip can swing an average. Look at the worst three clips by hand: if errors are names and numbers, a vocabulary or post-processing fix may help; if they are dropped phrases, the audio quality is the issue. Keep the reference transcripts, so you can re-score when a vendor ships a new model.

A model that scores 4 percent against another's 6 percent on your clips is a real difference only if the sample is large enough. With 15 one-minute clips you have roughly 2,000 reference words; treat gaps of a point or less as noise and decide on price, latency and features instead.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume