Run your own 10-clip word error test on Sume STT for 10 cents

Microsoft ranks MAI-Transcribe-2-Streaming first on Artificial Analysis. To know your own audio, score 10 one-minute clips with a 15-line WER function.

6 min readSume
All posts

Microsoft says MAI-Transcribe-2-Streaming ranks first on Artificial Analysis for the accuracy of both final and partial transcripts (read 2026-10-04). A public ranking is not your microphone, your accents or your product names. For about 10 cents you can measure your own: transcribe 10 one-minute clips with Sume STT at $0.01 per audio minute, compare against a human transcript, and compute word error rate. The scoring function below is 15 lines of Python.

Set up the clips

Pick 10 clips of 60 seconds that look like production: a quiet voice, a noisy one, a phone call, two speakers, a clip full of brand names. Write the correct transcript by hand for each. If the source is video, extract speech with audio detach first. Submit each to POST /v1/stt-1.0/transcribe with a public HTTPS audio_url and duration_seconds: 60 so the reservation matches the clip, and set language_code if you know it.

The scoring function

Word error rate is the edit distance between reference and hypothesis words, divided by the number of reference words. Lowercase and strip punctuation first, or the score punishes commas.

import re


def norm(text):
    return re.sub(r"[^\w\s']", "", text.lower()).split()


def wer(reference, hypothesis):
    ref, hyp = norm(reference), norm(hypothesis)
    prev = list(range(len(hyp) + 1))
    for i, r in enumerate(ref, 1):
        cur = [i]
        for j, h in enumerate(hyp, 1):
            cur.append(min(prev[j] + 1, cur[j - 1] + 1, prev[j - 1] + (r != h)))
        prev = cur
    return prev[-1] / max(len(ref), 1)

Reading the result

Called with the reference "Order the blue one, please." and the hypothesis "order the blue ones please", the function returns 0.2: five reference words, one substituted. Average over the 10 clips, then look at the worst two. A model with a better mean and a bad worst case may be worse for you.

What to log per clip (suggested columns, vendor facts read 2026-10-04)
ColumnWhy
Clip typeNoise and accent drive most of the spread
WERSingle comparable number
Brand-name missesCount separately; WER hides them
Seconds from request to resultSume jobs are async; streaming services quote partial latency instead
CostSume $0.01 per minute; MAI-Transcribe-2-Streaming $0.54 per hour on the introductory rate

Caveats

Streaming and file transcription are different products. Microsoft's headline latency is the first partial, just over 100 ms; a Sume job returns when the whole file is done, so compare accuracy on the final text, and treat latency as a separate question. Poll as described in jobs and results. Also remember that Sume fixes diarization and audio-event tagging server-side, so speaker labels are out of scope for this test.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume