Run your own 10-clip word error test on Sume STT for 10 cents
Microsoft ranks MAI-Transcribe-2-Streaming first on Artificial Analysis. To know your own audio, score 10 one-minute clips with a 15-line WER function.

Microsoft says MAI-Transcribe-2-Streaming ranks first on Artificial Analysis for the accuracy of both final and partial transcripts (read 2026-10-04). A public ranking is not your microphone, your accents or your product names. For about 10 cents you can measure your own: transcribe 10 one-minute clips with Sume STT at $0.01 per audio minute, compare against a human transcript, and compute word error rate. The scoring function below is 15 lines of Python.
Set up the clips
Pick 10 clips of 60 seconds that look like production: a quiet voice, a noisy one, a phone call, two speakers, a clip full of brand names. Write the correct transcript by hand for each. If the source is video, extract speech with audio detach first. Submit each to POST /v1/stt-1.0/transcribe with a public HTTPS audio_url and duration_seconds: 60 so the reservation matches the clip, and set language_code if you know it.
The scoring function
Word error rate is the edit distance between reference and hypothesis words, divided by the number of reference words. Lowercase and strip punctuation first, or the score punishes commas.
import re
def norm(text):
return re.sub(r"[^\w\s']", "", text.lower()).split()
def wer(reference, hypothesis):
ref, hyp = norm(reference), norm(hypothesis)
prev = list(range(len(hyp) + 1))
for i, r in enumerate(ref, 1):
cur = [i]
for j, h in enumerate(hyp, 1):
cur.append(min(prev[j] + 1, cur[j - 1] + 1, prev[j - 1] + (r != h)))
prev = cur
return prev[-1] / max(len(ref), 1)
Reading the result
Called with the reference "Order the blue one, please." and the hypothesis "order the blue ones please", the function returns 0.2: five reference words, one substituted. Average over the 10 clips, then look at the worst two. A model with a better mean and a bad worst case may be worse for you.
| Column | Why |
|---|---|
| Clip type | Noise and accent drive most of the spread |
| WER | Single comparable number |
| Brand-name misses | Count separately; WER hides them |
| Seconds from request to result | Sume jobs are async; streaming services quote partial latency instead |
| Cost | Sume $0.01 per minute; MAI-Transcribe-2-Streaming $0.54 per hour on the introductory rate |
Caveats
Streaming and file transcription are different products. Microsoft's headline latency is the first partial, just over 100 ms; a Sume job returns when the whole file is done, so compare accuracy on the final text, and treat latency as a separate question. Poll as described in jobs and results. Also remember that Sume fixes diarization and audio-event tagging server-side, so speaker labels are out of scope for this test.
Sources
Related posts
More in Developers
- Same prompt, four Sume video models: a Python script that logs cost
Submit one prompt to wan-3.0, minimax-h3, minimax-h3-max and seedance-2.5 on /v1/videos, poll each job, and print usage.cost per clip. Runnable as written.
- What to save from a Sume run when batch results expire at 30 days
OpenAI keeps batch output 30 days, Anthropic 29, Gemini 6 weeks. Which Sume run receipt fields to store so your records outlive any vendor retention window.
- Why a Sume scheduled run's events_url is null, and what to poll
Action runs always return events_url null; Format runs expose a phase timeline. What to poll for a scheduled run and how to read skipped.
- isTerminalJobStatus vs isTerminalRunStatus: skipped only ends runs
The Sume SDK has two terminal checks. Jobs end on completed, failed or canceled; runs also end on skipped. Reusing one for both breaks a custom poll loop.
Written by Sume