Test MAI-Transcribe-2-Streaming's accuracy claim on your own audio
Microsoft says its streaming transcriber ranks first on Artificial Analysis. A 10-clip test on your own audio, with a word error rate script and Sume STT.

Should you believe a first place on a public speech-to-text leaderboard for your own audio? Test it. A leaderboard averages other people's recordings; your calls, accents and product names may behave differently. This post gives a small test you can run on any transcriber, with a word error rate script that needs no API.
Microsoft's announcement (read 2026-10-04) says MAI-Transcribe-2-Streaming ranks first on Artificial Analysis for accuracy on both final and partial transcripts, produces first text just over 100 ms after receiving audio, covers 60 languages with continuous language detection, and costs $0.54 per hour of audio through year-end. Those are Microsoft's statements; none of them is a measurement of your audio.
Build the test set
Pick ten clips of 30 to 60 seconds that look like production: noisy ones, fast talkers, jargon, and at least one with a name your brand cares about. Type a reference transcript for each by hand. Do not start from a machine transcript, or you will grade the machine against itself.
- Keep punctuation and casing out of the comparison; compare lowercase words only.
- Write numbers the way you want them to appear, then decide up front whether
twentyand20count as a match. - Hold back two clips you never tune on.
Where Sume fits
Sume cannot run MAI-Transcribe-2-Streaming; its STT surface uses a provider the public contract keeps internal. What Sume can do is prepare audio and give you one more candidate to score. For a video source, audio detach returns a wav with channels: "mono" and sample_rate: 16000, the STT shape, at $0.01 per job. Then POST /v1/stt-1.0/transcribe returns text and words[] with start and end seconds.
Video inspect documents the STT rate as $0.01 per audio minute and says language_code is a hint you can omit for auto-detect. Confirm live prices in GET /v1/catalog before you budget.
Score it
Word error rate is substitutions plus deletions plus insertions, divided by the number of reference words. This script computes it for one pair of strings; run it per clip for each candidate and compare the totals.
import re
def words(s):
return re.findall(r"[a-z0-9']+", s.lower())
def wer(ref, hyp):
r, h = words(ref), words(hyp)
d = list(range(len(h) + 1))
for i in range(1, len(r) + 1):
prev, d[0] = d[0], i
for j in range(1, len(h) + 1):
cur = min(d[j] + 1, d[j - 1] + 1,
prev + (r[i - 1] != h[j - 1]))
prev, d[j] = d[j], cur
return d[len(h)] / max(len(r), 1)
print(round(wer("book a table for two at eight", "book a table for you at eight"), 3))
Read the result honestly
Ten short clips is a smoke test, not a benchmark. If two candidates land within a point or two of each other, the difference is noise, and price, language coverage and latency should decide. If one is far worse on your hard clips, that is the finding that matters, whatever the leaderboard says.
Streaming and file jobs also differ in shape: a streaming model returns partial text while audio arrives, and Sume STT is a job you poll. See live versus file transcription before you compare their prices.
Sources
Related posts
More in Developers
- Text to Dialogue continuity: 100-character context, 3 request IDs
ElevenLabs' Sept 28 changelog adds previous_text and future_text (100 chars max) and request-ID chaining (3 max). A Python limit check.
- Thin vs full webhook payloads: how Sume's job events work
Stripe made thin events generally available for API v1. Sume's job webhooks carry the result in the payload, are keyed by job_id, and can be redelivered.
- Three Sume throttle signals: rate_limited, queue_full, 503
OpenAI now splits 429 (traffic rising too fast) from 503 (overload). Sume has three: 429 rate_limited, 429 queue_full and 503 provider_capacity_exceeded.
- Track AI video spend per job: Synthesia Billing API vs Sume job cost
Synthesia added a Billing API and auto top-up. On Sume, every finished job reports usage.cost, so you can keep a per-job ledger without a billing endpoint.
Written by Sume