Measure word error rate on your own clips with Python and Sume STT

Microsoft charts 2.5% word error rate for MAI-Transcribe-2-Streaming. Measure yours: Sume STT plus a 20-line word-level edit distance in plain Python.

5 min readSume
All posts

Word error rate is (substitutions + deletions + insertions) divided by the number of words in a human reference. A word-level edit distance in about 20 lines of standard-library Python computes it, and Sume STT supplies the machine transcript. Microsoft's MAI-Transcribe-2 page charts 2.5% final word error rate for MAI-Transcribe-2-Streaming, from a source dated 28 September 2026. That number is on someone else's test set, so measure your own clips.

I read Microsoft's MAI-Transcribe-2 page, the Sume job docs and the Sume API reference on 2026-10-03, and ran the code below against a mocked API. I did not run it against live audio.

What does Microsoft's number measure?

The page attributes its streaming chart to the Artificial Analysis Speech to Text (Streaming) leaderboard dated 28 September 2026, and its multilingual chart to the FLEURS test set. A vendor figure is a score on that set, with that normalization, in that language mix. Your accents, phone audio and product names are different, so the figure is a reason to test, not a forecast. The page also lists the streaming model at #1 on the Artificial Analysis accuracy leaderboard, which is again a ranking on their benchmark.

How do you get a reference and a hypothesis?

Pick 10 to 20 clips of 10 to 60 seconds that look like your real traffic. Type what is actually said into reference.txt, exactly, including repeated words. Then transcribe each clip. Sume STT is an asynchronous job: poll status_url until terminal is true, then read result_url, where data.result.text holds the transcript. STT lists $0.01 per audio minute.

What does the code look like?

The script normalises case, hyphens and punctuation before counting, so Hello, World! equals hello world. It keeps apostrophes. Set SUME_API_KEY and CLIP_URL, and put the human text in reference.txt.

import json, os, re, time, urllib.request

def api(method, url, body=None):
    req = urllib.request.Request(url, method=method, data=body and json.dumps(body).encode(),
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"})
    with urllib.request.urlopen(req) as r:
        return json.load(r)["data"]

def transcribe(audio_url):
    job = api("POST", "https://api.sume.com/v1/stt-1.0/transcribe", {"audio_url": audio_url, "mode": "async"})
    while not api("GET", job["status_url"])["terminal"]:
        time.sleep(2)
    return api("GET", job["result_url"])["result"]["text"]

def words(text):
    return re.sub(r"[^\w\s']", "", text.lower().replace("-", " ")).split()

def wer(reference, hypothesis):
    ref, hyp = words(reference), words(hypothesis)
    prev = list(range(len(hyp) + 1))
    for i, r in enumerate(ref, 1):
        cur = [i]
        for j, h in enumerate(hyp, 1):
            cur.append(min(prev[j] + 1, cur[j - 1] + 1, prev[j - 1] + (r != h)))
        prev = cur
    return prev[-1] / max(len(ref), 1)

if __name__ == "__main__":
    reference = open("reference.txt").read()
    print(f"WER {wer(reference, transcribe(os.environ['CLIP_URL'])):.1%}")

How should you read the result?

Sum the edits and the reference words over all clips instead of averaging per-clip percentages, so one 3-word clip does not weigh as much as a 60-second one. Normalization changes the score: spelling out numbers or keeping punctuation moves it, so fix your rules before comparing vendors.

Word error rate worked example, read 2026-10-03
ReferenceMachine textEditsWER
Hello, world! This is a test of the caption pipeline.hello world this is test of the caption pipe line1 deletion, 1 substitution, 1 insertion3 / 10 = 30.0%

Sources

Related posts

More in Developers

All Developers posts

Written by Sume