MAI-Transcribe-2-Streaming lists no word timestamps; use batch

Microsoft's page shows word timestamps, diarization and keyword biasing for batch MAI-Transcribe-2 but not Streaming. Sume STT returns words. Read 2026-10-08.

5 min readSume
All posts

If you want captions that highlight words, do not pick MAI-Transcribe-2-Streaming: Microsoft's model page shows word-level timestamps, diarization, contextual biasing and clean/verbatim style as supported for batch MAI-Transcribe-2 and as "NO" for the streaming model. The streaming model is for live audio, with about 120 ms to a first partial transcript. Sume STT is a file job that returns word timings and optional sentence segments, so it sits on the batch side of that split.

What the page's feature rows say

The table is transcribed from the model page. Prices are the page's introductory figures.

MAI transcription models, read 2026-10-08
ModelLanguagesReal-timeLatencyWord timestampsKeyword biasingDiarizationClean/verbatim$ per audio hour
MAI-Transcribe-2 (batch)60No~1 hr audio in ~10 sYesYesYesYes$0.10
MAI-Transcribe-2-Streaming60Yes~120 ms first partial; ~128 ms end-to-endNoNoNoNo$0.54
MAI-Transcribe-1.543No~1 hr audio in ~20 sNoNoNoNo$0.36

Where Sume STT sits

Sume STT 1.0 costs $0.01 per audio minute ($0.60 per hour) and is used through POST /v1/stt-1.0/transcribe, with a job you poll or receive by webhook. The transcript shape documented for video inspect is text, words[] and optional sentence segments[]; a language_code hint is optional and auto-detect is the default (Video inspect). A request is capped at 600 seconds. I did not find a documented streaming mode, diarization option or keyword-biasing option in the docs I read, so do not assume Sume STT has them.

For captions on a finished video, the comparison is therefore batch against batch: $0.10 an hour from Microsoft's page against $0.60 an hour at Sume. Pay Sume's price only if you need the transcript to land inside the Sume job flow.

A split pipeline

Some teams run both kinds of recognizer: a streaming model for the live preview and a batch model for the final caption pass, because the final pass needs timestamps and the preview needs speed. On the page's prices, one hour of audio through both Microsoft models is $0.54 + $0.10 = $0.64, which is about the same as one hour of Sume STT at $0.60. That is only a cost comparison; accuracy needs your own test files.

Reading the page row by row

Three details on the page change the decision. First, the streaming row shows a time to first partial of about 120 ms and an end-to-end latency of about 128 ms, while the batch row shows no first-partial figure and instead states that an hour of audio is transcribed in about 10 seconds. Second, only the batch row lists a transcription style choice between clean and verbatim, which matters when a caption editor needs the filler words removed or kept. Third, the older MAI-Transcribe-1.5 row shows 43 languages at $0.36 an hour with none of the optional features, so moving up to the 2-series adds 17 languages and, for the batch model, the feature set, at a lower hourly price.

All of these are the page's own fields as read on 2026-10-08. The page marks the 2-series prices introductory and does not state when they change, so a budget that runs past the end of the year should include a margin for that.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume