MAI-Transcribe-2-Streaming lists no word timestamps; use batch
Microsoft's page shows word timestamps, diarization and keyword biasing for batch MAI-Transcribe-2 but not Streaming. Sume STT returns words. Read 2026-10-08.

If you want captions that highlight words, do not pick MAI-Transcribe-2-Streaming: Microsoft's model page shows word-level timestamps, diarization, contextual biasing and clean/verbatim style as supported for batch MAI-Transcribe-2 and as "NO" for the streaming model. The streaming model is for live audio, with about 120 ms to a first partial transcript. Sume STT is a file job that returns word timings and optional sentence segments, so it sits on the batch side of that split.
What the page's feature rows say
The table is transcribed from the model page. Prices are the page's introductory figures.
| Model | Languages | Real-time | Latency | Word timestamps | Keyword biasing | Diarization | Clean/verbatim | $ per audio hour |
|---|---|---|---|---|---|---|---|---|
| MAI-Transcribe-2 (batch) | 60 | No | ~1 hr audio in ~10 s | Yes | Yes | Yes | Yes | $0.10 |
| MAI-Transcribe-2-Streaming | 60 | Yes | ~120 ms first partial; ~128 ms end-to-end | No | No | No | No | $0.54 |
| MAI-Transcribe-1.5 | 43 | No | ~1 hr audio in ~20 s | No | No | No | No | $0.36 |
Where Sume STT sits
Sume STT 1.0 costs $0.01 per audio minute ($0.60 per hour) and is used through POST /v1/stt-1.0/transcribe, with a job you poll or receive by webhook. The transcript shape documented for video inspect is text, words[] and optional sentence segments[]; a language_code hint is optional and auto-detect is the default (Video inspect). A request is capped at 600 seconds. I did not find a documented streaming mode, diarization option or keyword-biasing option in the docs I read, so do not assume Sume STT has them.
For captions on a finished video, the comparison is therefore batch against batch: $0.10 an hour from Microsoft's page against $0.60 an hour at Sume. Pay Sume's price only if you need the transcript to land inside the Sume job flow.
A split pipeline
Some teams run both kinds of recognizer: a streaming model for the live preview and a batch model for the final caption pass, because the final pass needs timestamps and the preview needs speed. On the page's prices, one hour of audio through both Microsoft models is $0.54 + $0.10 = $0.64, which is about the same as one hour of Sume STT at $0.60. That is only a cost comparison; accuracy needs your own test files.
Reading the page row by row
Three details on the page change the decision. First, the streaming row shows a time to first partial of about 120 ms and an end-to-end latency of about 128 ms, while the batch row shows no first-partial figure and instead states that an hour of audio is transcribed in about 10 seconds. Second, only the batch row lists a transcription style choice between clean and verbatim, which matters when a caption editor needs the filler words removed or kept. Third, the older MAI-Transcribe-1.5 row shows 43 languages at $0.36 an hour with none of the optional features, so moving up to the 2-series adds 17 languages and, for the batch model, the feature set, at a lower hourly price.
All of these are the page's own fields as read on 2026-10-08. The page marks the 2-series prices introductory and does not state when they change, so a budget that runs past the end of the year should include a margin for that.
Sources
Related posts
More in Comparisons
- MAI-Transcribe-2-Streaming is a preview with no SLA: add a fallback
Microsoft marks MAI-Transcribe-2-Streaming public preview with no SLA. Record call audio and keep Sume STT as a recorded fallback at about one cent per minute.
- MAI-Transcribe-2-Streaming vs Sume STT: 100 hours, $54 vs $60
A tracker lists MAI-Transcribe-2-Streaming at $0.54 per hour (introductory through 2026) and Sume STT is $0.01 per audio minute, or $0.60 per hour.
- MAI-Voice-2.1 and MAI-Transcribe-2: what Sume lists for voice and STT
Microsoft AI's news page names MAI-Voice-2.1 and MAI-Transcribe-2. Sume lists neither; it lists Sonic TTS (38 micros a character) and STT at $0.01 a minute.
- MAI-Voice-2.1 or Sume TTS for a presenter ad read: 3 prices
For a 700-character ad read, MAI-Voice-2.1 is about 1.5 cents, Flash 1 cent and Sume TTS 4 cents. Fabric needs the audio as a public URL.
Written by Sume