Test a 'noisy audio' transcription claim: 5 clips for 5 cents on Sume

Microsoft says MAI-Transcribe-2 handles noisy audio. Run five 60-second noisy clips through Sume STT for $0.05 and judge the text yourself.

5 min readSume
All posts

Run five 60-second clips of your own noisiest audio through Sume STT 1.0: five minutes at $0.01 per audio minute is $0.05, and the same five minutes on MAI-Transcribe-2-Streaming at $0.54 an hour is about $0.045. Microsoft's MAI-Transcribe-2 page says the model turns noisy audio into precise, domain-specific transcripts, and no vendor claim about your audio replaces a test on your audio.

Rates are from Microsoft's announcement and the Sume API reference, read 2026-10-05. The clip count and length are the test design suggested here, not a benchmark.

What the test costs

Five one-minute clips cost cents on either service, so the limit is your time to listen, not the bill.

Cost of five 60-second clips (5 audio minutes) at list rates (read 2026-10-05)
ServiceRateCost of 5 minutesNotes
MAI-Transcribe-2-Streaming$0.54 per audio hour (intro rate through end of 2026)$0.045Streaming model; Microsoft names Artificial Analysis and FLEURS scores on its page
Sume STT 1.0$0.01 per audio minute$0.05One job per clip, duration_seconds 60

How to build a fair noisy-audio test

Pick clips that look like your production audio, not like a demo. Keep the clips fixed so you can re-run the test when a model changes.

  • Choose five real clips: a street interview, a kitchen, a car, a crowded room, a poor phone line.
  • Write a reference transcript for each by hand before looking at any machine output.
  • Send each clip with language_code set when you know the language; omit it to test auto-detect.
  • Read words[] timings as well as text: a transcript can have the right words and the wrong times.
  • Compare the two versions side by side and mark each wrong word; do not average a score across clips of very different quality.

Run it on Sume

This submits five hosted clips and prints each transcript with the language the service reports. The audio_url must be public HTTPS; Sume prefers its own media host.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
    r = requests.post(API + path, headers={**H, "Idempotency-Key": key}, json=body, timeout=60)
    r.raise_for_status()
    job = r.json()["data"]
    while True:
        s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
        if s["terminal"]:
            break
        time.sleep(2)
    if s["sume_status"] != "completed":
        raise RuntimeError(s["sume_status"])
    return requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]

clips = os.environ["NOISY_CLIP_URLS"].split(",")
for i, url in enumerate(clips, 1):
    body = {"audio_url": url, "duration_seconds": 60}
    out = run("/v1/stt-1.0/transcribe", body, f"noisy-test-{i}")
    print(i, out.get("language_code"), out["text"][:200])

What the result can and cannot tell you

Five clips is a sanity check, not a measurement. If one service is clearly worse on your worst clip, that is a useful signal. If they look alike, widen the sample. Microsoft's accuracy claims are for its own benchmarks; Sume's STT is billed per audio minute and runs a provider model behind a Sume job, so a test on your clips is the only comparison that matches your use.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume