MAI-Transcribe-2-Streaming has no timestamps: time the captions

MAI-Transcribe-2-Streaming lists no timestamps. For timed captions, get word times from Sume STT or burn your text with script_text.

5 min readSume
All posts

MAI-Transcribe-2-Streaming returns text but no word timestamps, diarization, style options or biasing, according to Microsoft's model page. To make timed captions from its transcript, you need times from somewhere else. Sume STT returns word timings on every job, and Sume video captions can align your text to those timings through script_text (read 2026-10-07).

What the streaming model gives you

The model page lists 2-Streaming at roughly 120 ms to a first partial and 128 ms end to end, ranked first on Artificial Analysis, with an introductory price of $0.54 per hour. The listed feature set is shorter than the batch model's. MAI-Transcribe-2 has word timestamps, biasing, diarization and clean or verbatim output. The streaming model has none of these.

The Realtime API page describes the events: delta carries final text to append, intermediate replaces a provisional suffix, and committed and completed mark the turn. The page documents no start or end times on them.

What each MAI-Transcribe model lists (read 2026-10-07)
Feature22-Streaming
Word timestampslistednot listed
Diarizationlistednot listed
Biasinglistednot listed
Price per hour (introductory)$0.10$0.54

Two ways to time it on Sume

The first is to take the same audio, upload it to your workspace, and run Sume STT at $0.01 per audio minute. The result has words with start and end times, so you can build captions from it directly. Use STT's own text if it is good enough, or swap in the streaming transcript where it is better.

The second is for a finished video. Send the video to Sume video captions with the MAI text in script_text. Sume keeps the speech-to-text timings as the source of truth for time and aligns your text to them. It costs $0.20 for a video up to 60 seconds. If the texts differ too much, the job fails with script_alignment_mismatch, and you correct the script and retry.

The catch

Alignment works when the two transcripts say mostly the same words. If MAI heard a proper noun differently from Sume STT, the aligner has to bridge the gap. Read the first run, fix the offending words in the script text, and send it again. You pay again for each render.

If the video has no audible speech, captions from script_text fail with caption_no_speech. In that case send cues with text, start and end, and Sume burns them without speech recognition.

  • Streaming transcript plus Sume STT timing: two vendors, one caption file.
  • script_text: one render, $0.20, aligned to Sume's own timings.
  • Silent clip: use cues, not script_text.

Is the second vendor worth it?

If you only need captions for recorded video, Sume alone may be enough: video captions can run speech-to-text for you and burn the result, with no MAI step. MAI-Transcribe-2-Streaming earns its place when the audio is live and you need text in about a tenth of a second. Recorded captions do not need that speed.

Use the streaming model for the live view, then send the saved recording through Sume for the final, timed captions.

A practical sequence

Store the audio as you stream it, in the PCM16 mono format the streaming API takes. When the session ends, wrap it in a WAV header, upload it to your workspace, and run STT with the language hint if you know it. Compare the final MAI text with the Sume STT text on the first few minutes, and decide which transcript is the script.

Then caption the video. One 60-second render is $0.20, and a ten-minute video is ten renders, so $2.00 before any retries.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume