MAI-Transcribe-2-Streaming 2x faster: what it means for a file
Microsoft says words appear 2x faster than its closest competitor. That is a live-audio claim. For a recorded file on Sume, measure submit-to-result yourself.

Microsoft's launch post for MAI-Transcribe-2-Streaming says first hypotheses, which it calls partials, arrive in just over 100 ms, and that words appear 2x faster than the closest competitor. Both numbers describe a stream: audio goes in while someone is still talking and text comes out. A recorded file has no speaker waiting, so the claim does not transfer to it, and Sume's STT 1.0 is a batch job, not a stream.
What to measure on a file
Sume's route takes a public HTTPS audio_url and returns a job. The result carries text, words[] timings and, if you ask, sentence segments. The wait is bounded by wait_timeout_seconds, with a maximum of 30, and a longer job must be polled from status_url. So the number to measure is the time from submit to result_ready.
A fair test
Time three clips of different lengths, for example one short, one near a minute and one near the 10-minute cap, and record the elapsed seconds in your own table. Do not carry a figure from this page into a plan, because nothing here is a benchmark. Send duration_seconds so the reservation matches the file.
start=$(date +%s)
curl -s -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-timing-001" \
-d '{"audio_url": "'"$AUDIO_URL"'", "duration_seconds": 60, "mode": "sync", "wait_timeout_seconds": 30}' > out.json
echo "elapsed: $(( $(date +%s) - start )) s"; head -c 400 out.jsonThe claims side by side, read 2026-10-06:
| Item | MAI-Transcribe-2-Streaming (vendor) | Sume STT 1.0 (docs) |
|---|---|---|
| Mode | Streaming | Batch job |
| Languages | 60, continuous detection | Optional language_code, auto-detect if omitted |
| Speed claim | Partials in just over 100 ms | Measure it yourself |
| Max input | Not stated | 10 minutes per file |
| Price | $0.54 per hour through the end of the year | $0.01 per audio minute |
Reading a vendor speed claim
A latency figure always has a condition. Microsoft's post ties it to partials from a stream and a comparison with a named kind of competitor, and gives no figure for a finished file. A fair way to use it is as a reason to test the model for live captions, not as a number to copy into a batch plan.
For a recorded file, the cost of waiting is usually less important than the cost of an error. Check accuracy on your own audio first, then time it. The job is bounded at 30 seconds of waiting in sync mode, and longer files simply continue in the background until you read result.
If you need captions while a person speaks, a streaming model is the right tool. If you have the file already, a batch job gives you word timings you can reuse for captions. The Sume rate is the public rate from the repository fixtures, so confirm it in GET /v1/catalog.
Sources
Related posts
More in Comparisons
- MAI-Voice-2.1 clones from seconds: how Sume clones from a clip
Microsoft says MAI-Voice-2.1 clones from a few seconds of audio with consent guardrails. In Sume Agents, voices_create clones from an https audio clip.
- MAI-Voice-2.1 on Foundry and Vercel; Sume's TTS Router is Sonic only
Microsoft lists Foundry, MAI Playground, Vercel and OpenRouter for MAI-Voice-2.1. Sume's TTS Router lists only Cartesia Sonic ids, so check the catalog.
- MAI-Voice-2.1 or Sume TTS? Three questions that decide it
Live speech, extra outputs or lowest price per character? A short guide to MAI-Voice-2.1 and Flash against the Sume TTS Router, with 10M-character math.
- Shortest AI video clip by model: MiniMax H3 4 s, Sume row 5 s
MiniMax lists 4 to 15 seconds for H3. Sume's row starts at 5. Here are the minimum durations in the catalog and how to trim a longer clip.
Written by Sume