Measure word error rate on your own clips with Python and Sume STT
Microsoft charts 2.5% word error rate for MAI-Transcribe-2-Streaming. Measure yours: Sume STT plus a 20-line word-level edit distance in plain Python.

Word error rate is (substitutions + deletions + insertions) divided by the number of words in a human reference. A word-level edit distance in about 20 lines of standard-library Python computes it, and Sume STT supplies the machine transcript. Microsoft's MAI-Transcribe-2 page charts 2.5% final word error rate for MAI-Transcribe-2-Streaming, from a source dated 28 September 2026. That number is on someone else's test set, so measure your own clips.
I read Microsoft's MAI-Transcribe-2 page, the Sume job docs and the Sume API reference on 2026-10-03, and ran the code below against a mocked API. I did not run it against live audio.
What does Microsoft's number measure?
The page attributes its streaming chart to the Artificial Analysis Speech to Text (Streaming) leaderboard dated 28 September 2026, and its multilingual chart to the FLEURS test set. A vendor figure is a score on that set, with that normalization, in that language mix. Your accents, phone audio and product names are different, so the figure is a reason to test, not a forecast. The page also lists the streaming model at #1 on the Artificial Analysis accuracy leaderboard, which is again a ranking on their benchmark.
How do you get a reference and a hypothesis?
Pick 10 to 20 clips of 10 to 60 seconds that look like your real traffic. Type what is actually said into reference.txt, exactly, including repeated words. Then transcribe each clip. Sume STT is an asynchronous job: poll status_url until terminal is true, then read result_url, where data.result.text holds the transcript. STT lists $0.01 per audio minute.
What does the code look like?
The script normalises case, hyphens and punctuation before counting, so Hello, World! equals hello world. It keeps apostrophes. Set SUME_API_KEY and CLIP_URL, and put the human text in reference.txt.
import json, os, re, time, urllib.request
def api(method, url, body=None):
req = urllib.request.Request(url, method=method, data=body and json.dumps(body).encode(),
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"], "Content-Type": "application/json"})
with urllib.request.urlopen(req) as r:
return json.load(r)["data"]
def transcribe(audio_url):
job = api("POST", "https://api.sume.com/v1/stt-1.0/transcribe", {"audio_url": audio_url, "mode": "async"})
while not api("GET", job["status_url"])["terminal"]:
time.sleep(2)
return api("GET", job["result_url"])["result"]["text"]
def words(text):
return re.sub(r"[^\w\s']", "", text.lower().replace("-", " ")).split()
def wer(reference, hypothesis):
ref, hyp = words(reference), words(hypothesis)
prev = list(range(len(hyp) + 1))
for i, r in enumerate(ref, 1):
cur = [i]
for j, h in enumerate(hyp, 1):
cur.append(min(prev[j] + 1, cur[j - 1] + 1, prev[j - 1] + (r != h)))
prev = cur
return prev[-1] / max(len(ref), 1)
if __name__ == "__main__":
reference = open("reference.txt").read()
print(f"WER {wer(reference, transcribe(os.environ['CLIP_URL'])):.1%}")How should you read the result?
Sum the edits and the reference words over all clips instead of averaging per-clip percentages, so one 3-word clip does not weigh as much as a 60-second one. Normalization changes the score: spelling out numbers or keeping punctuation moves it, so fix your rules before comparing vendors.
| Reference | Machine text | Edits | WER |
|---|---|---|---|
| Hello, world! This is a test of the caption pipeline. | hello world this is test of the caption pipe line | 1 deletion, 1 substitution, 1 insertion | 3 / 10 = 30.0% |
Sources
Related posts
More in Developers
- MAI-Transcribe-2-Streaming '2x faster': measure your caption delay
Microsoft says words appear 2x faster than its closest competitor. Measure your own delay from speech to caption with a p50 and p95 script before you switch.
- Microsoft Agent Framework hosted MCP tool for Sume (Python)
Attach https://mcp.sume.com/mcp to an Agent Framework agent with get_mcp_tool, allowed_tools and approval_mode. Foundry runs the calls, so mind the key.
- Migrate Video 1.0 to sume/auto on POST /v1/videos, field by field
Video 1.0 is retiring soon. Which fields move, which are ignored or rejected, and the request to send on POST /v1/videos with sume/auto.
- Migrate POST /v1/video-router/generate to /v1/videos: a field map
Same model ids, same jobs, different wire. How image_url and reference_image_urls become frame_images and input_references, and what stays on Video Router.
Written by Sume