Test a 'noisy audio' transcription claim: 5 clips for 5 cents on Sume
Microsoft says MAI-Transcribe-2 handles noisy audio. Run five 60-second noisy clips through Sume STT for $0.05 and judge the text yourself.

Run five 60-second clips of your own noisiest audio through Sume STT 1.0: five minutes at $0.01 per audio minute is $0.05, and the same five minutes on MAI-Transcribe-2-Streaming at $0.54 an hour is about $0.045. Microsoft's MAI-Transcribe-2 page says the model turns noisy audio into precise, domain-specific transcripts, and no vendor claim about your audio replaces a test on your audio.
Rates are from Microsoft's announcement and the Sume API reference, read 2026-10-05. The clip count and length are the test design suggested here, not a benchmark.
What the test costs
Five one-minute clips cost cents on either service, so the limit is your time to listen, not the bill.
| Service | Rate | Cost of 5 minutes | Notes |
|---|---|---|---|
| MAI-Transcribe-2-Streaming | $0.54 per audio hour (intro rate through end of 2026) | $0.045 | Streaming model; Microsoft names Artificial Analysis and FLEURS scores on its page |
| Sume STT 1.0 | $0.01 per audio minute | $0.05 | One job per clip, duration_seconds 60 |
How to build a fair noisy-audio test
Pick clips that look like your production audio, not like a demo. Keep the clips fixed so you can re-run the test when a model changes.
- Choose five real clips: a street interview, a kitchen, a car, a crowded room, a poor phone line.
- Write a reference transcript for each by hand before looking at any machine output.
- Send each clip with
language_codeset when you know the language; omit it to test auto-detect. - Read
words[]timings as well astext: a transcript can have the right words and the wrong times. - Compare the two versions side by side and mark each wrong word; do not average a score across clips of very different quality.
Run it on Sume
This submits five hosted clips and prints each transcript with the language the service reports. The audio_url must be public HTTPS; Sume prefers its own media host.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
r.raise_for_status()
job = r.json()["data"]
while True:
s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
if s["terminal"]:
break
time.sleep(2)
if s["sume_status"] != "completed":
raise RuntimeError(s["sume_status"])
return requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
clips = os.environ["NOISY_CLIP_URLS"].split(",")
for i, url in enumerate(clips, 1):
body = {"audio_url": url, "duration_seconds": 60}
out = run("/v1/stt-1.0/transcribe", body, f"noisy-test-{i}")
print(i, out.get("language_code"), out["text"][:200])
What the result can and cannot tell you
Five clips is a sanity check, not a measurement. If one service is clearly worse on your worst clip, that is a useful signal. If they look alike, widen the sample. Microsoft's accuracy claims are for its own benchmarks; Sume's STT is billed per audio minute and runs a provider model behind a Sume job, so a test on your clips is the only comparison that matches your use.
Sources
Related posts
More in Comparisons
- Text-to-speech price per million characters: 11 rates vs Sume
Eleven TTS rates converted to dollars per 1M characters, from $11 to $150. Sume TTS 1.0 is $47.50. Each is read from the vendor page, and caveats are listed.
- 3-hour hearing transcript: MAI Streaming vs Sume STT speaker labels
A 180-minute hearing costs about $1.62 on MAI-Transcribe-2-Streaming and $1.81 on Sume STT with the split, but Sume's result has no speaker labels.
- Wan 3.0, H3 and H3 Max tie at $0.0625 a second at 480p on Sume
At 480p, wan-3.0, minimax-h3 and minimax-h3-max all bill $0.0625 a second on Sume, $0.63 for 10 seconds. Break the tie on length and tiers above.
- Does TikTok auto-label AI video? Only its own AI effects
TikTok labels content made with its official AI effects automatically. A Sume clip is made outside the app, so you disclose it yourself when it looks real.
Written by Sume