MAI-Transcribe-2-Streaming or a batch STT job: which fits?
MAI-Transcribe-2-Streaming returns first partials in just over 100 ms. Sume's speech-to-text is a batch job up to 10 minutes at $0.01 per minute.

Use a streaming model such as MAI-Transcribe-2-Streaming when text must appear while someone is still speaking. Use a batch job like Sume's speech-to-text when you have a finished recording and want a transcript with sentence timings. Microsoft reports first partials in just over 100 ms and 60 languages for MAI-Transcribe-2-Streaming.
Streaming versus batch
| Item | MAI-Transcribe-2-Streaming | Sume STT |
|---|---|---|
| Mode | Streaming | Job on a public audio_url |
| Price | $0.54 per hour through year end | $0.01 per audio minute ($0.60 per hour) |
| Latency claim | First partials just over 100 ms | Not claimed |
| Max input | Continuous stream | 10 minutes (600 s) per job |
| Output | Partial and final text | Sentence segmentation |
Reported accuracy
Unite.ai reports 2.5% final word error rate, 2.8% for the first partial and 0.13 seconds to the final transcript. Those are vendor-reported; I ran no test.
Batch with Sume
Send audio_url, an optional language_code and duration_seconds from 1 to 600. Sentence segmentation takes boundary_lead_ms from 0 to 500 with a default of 70.
- diarize and tag_audio_events are fixed on the server and a request that sends them returns 400.
- Longer audio: split with Timeline audio first, $0.01 per job.
- Check the video has an audio track before transcribing.
Where each is available
MAI-Transcribe-2-Streaming is listed on Foundry, MAI Playground, Vercel, OpenRouter and Azure Voice Live, with LiveKit coming soon.
Sources
Related posts
More in Models
- Can I use MAI-Voice-2.1 audio commercially? What Microsoft says
Microsoft's MAI voices page says it holds full licensing rights for commercial use, the service is in public preview with no SLA, and cloning is gated.
- MAI-Voice-2.1 emotion control vs Sume's emotion field
Microsoft lists emotion control on MAI-Voice-2.1. Sume's TTS takes a free-text emotion string, speed and volume. What each gives you, and how to test it.
- MAI-Voice-2.1 English voices with call-centre or audiobook styles
customer_call_center and audiobook styles sit on en-US Grant, Harper, Sage and en-GB Emily, Harry. Iris, Jasper, Dhruv and Priya are neutral only.
- Which Korean voices does MAI-Voice-2.1 have, and with what styles?
MAI-Voice-2.1 lists four ko-KR voices: Grant, Harper (role styles), Haena (14 emotion styles) and Junho (12). Pair a voice with the right caption style.
Written by Sume