MAI-Transcribe-2-Streaming 2x faster: what it means for a file

Microsoft says words appear 2x faster than its closest competitor. That is a live-audio claim. For a recorded file on Sume, measure submit-to-result yourself.

4 min readSume
All posts

Microsoft's launch post for MAI-Transcribe-2-Streaming says first hypotheses, which it calls partials, arrive in just over 100 ms, and that words appear 2x faster than the closest competitor. Both numbers describe a stream: audio goes in while someone is still talking and text comes out. A recorded file has no speaker waiting, so the claim does not transfer to it, and Sume's STT 1.0 is a batch job, not a stream.

What to measure on a file

Sume's route takes a public HTTPS audio_url and returns a job. The result carries text, words[] timings and, if you ask, sentence segments. The wait is bounded by wait_timeout_seconds, with a maximum of 30, and a longer job must be polled from status_url. So the number to measure is the time from submit to result_ready.

A fair test

Time three clips of different lengths, for example one short, one near a minute and one near the 10-minute cap, and record the elapsed seconds in your own table. Do not carry a figure from this page into a plan, because nothing here is a benchmark. Send duration_seconds so the reservation matches the file.

start=$(date +%s)
curl -s -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stt-timing-001" \
  -d '{"audio_url": "'"$AUDIO_URL"'", "duration_seconds": 60, "mode": "sync", "wait_timeout_seconds": 30}' > out.json
echo "elapsed: $(( $(date +%s) - start )) s"; head -c 400 out.json

The claims side by side, read 2026-10-06:

Streaming claim versus file job, read 2026-10-06
ItemMAI-Transcribe-2-Streaming (vendor)Sume STT 1.0 (docs)
ModeStreamingBatch job
Languages60, continuous detectionOptional language_code, auto-detect if omitted
Speed claimPartials in just over 100 msMeasure it yourself
Max inputNot stated10 minutes per file
Price$0.54 per hour through the end of the year$0.01 per audio minute

Reading a vendor speed claim

A latency figure always has a condition. Microsoft's post ties it to partials from a stream and a comparison with a named kind of competitor, and gives no figure for a finished file. A fair way to use it is as a reason to test the model for live captions, not as a number to copy into a batch plan.

For a recorded file, the cost of waiting is usually less important than the cost of an error. Check accuracy on your own audio first, then time it. The job is bounded at 30 seconds of waiting in sync mode, and longer files simply continue in the background until you read result.

If you need captions while a person speaks, a streaming model is the right tool. If you have the file already, a batch job gives you word timings you can reuse for captions. The Sume rate is the public rate from the repository fixtures, so confirm it in GET /v1/catalog.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume