MAI-Transcribe-2-Streaming has no timestamps: time the captions
MAI-Transcribe-2-Streaming lists no timestamps. For timed captions, get word times from Sume STT or burn your text with script_text.

MAI-Transcribe-2-Streaming returns text but no word timestamps, diarization, style options or biasing, according to Microsoft's model page. To make timed captions from its transcript, you need times from somewhere else. Sume STT returns word timings on every job, and Sume video captions can align your text to those timings through script_text (read 2026-10-07).
What the streaming model gives you
The model page lists 2-Streaming at roughly 120 ms to a first partial and 128 ms end to end, ranked first on Artificial Analysis, with an introductory price of $0.54 per hour. The listed feature set is shorter than the batch model's. MAI-Transcribe-2 has word timestamps, biasing, diarization and clean or verbatim output. The streaming model has none of these.
The Realtime API page describes the events: delta carries final text to append, intermediate replaces a provisional suffix, and committed and completed mark the turn. The page documents no start or end times on them.
| Feature | 2 | 2-Streaming |
|---|---|---|
| Word timestamps | listed | not listed |
| Diarization | listed | not listed |
| Biasing | listed | not listed |
| Price per hour (introductory) | $0.10 | $0.54 |
Two ways to time it on Sume
The first is to take the same audio, upload it to your workspace, and run Sume STT at $0.01 per audio minute. The result has words with start and end times, so you can build captions from it directly. Use STT's own text if it is good enough, or swap in the streaming transcript where it is better.
The second is for a finished video. Send the video to Sume video captions with the MAI text in script_text. Sume keeps the speech-to-text timings as the source of truth for time and aligns your text to them. It costs $0.20 for a video up to 60 seconds. If the texts differ too much, the job fails with script_alignment_mismatch, and you correct the script and retry.
The catch
Alignment works when the two transcripts say mostly the same words. If MAI heard a proper noun differently from Sume STT, the aligner has to bridge the gap. Read the first run, fix the offending words in the script text, and send it again. You pay again for each render.
If the video has no audible speech, captions from script_text fail with caption_no_speech. In that case send cues with text, start and end, and Sume burns them without speech recognition.
- Streaming transcript plus Sume STT timing: two vendors, one caption file.
- script_text: one render, $0.20, aligned to Sume's own timings.
- Silent clip: use cues, not script_text.
Is the second vendor worth it?
If you only need captions for recorded video, Sume alone may be enough: video captions can run speech-to-text for you and burn the result, with no MAI step. MAI-Transcribe-2-Streaming earns its place when the audio is live and you need text in about a tenth of a second. Recorded captions do not need that speed.
Use the streaming model for the live view, then send the saved recording through Sume for the final, timed captions.
A practical sequence
Store the audio as you stream it, in the PCM16 mono format the streaming API takes. When the session ends, wrap it in a WAV header, upload it to your workspace, and run STT with the language hint if you know it. Compare the final MAI text with the Sume STT text on the first few minutes, and decide which transcript is the script.
Then caption the video. One 60-second render is $0.20, and a ten-minute video is ten renders, so $2.00 before any retries.
Sources
Related posts
More in Developers
- Mandarin Chinese speech to text API: Sume STT language_code zh
Transcribe Mandarin audio with Sume STT using language_code zh, then check the result and timings. $0.01 per audio minute and no accuracy claim without a test.
- mask_url on ChatGPT Image 2: not listed, only the 2.5 rows take it
Sume lists mask_url only on ChatGPT Image 2.5 Flare and Sunburst. Sending it to ChatGPT Image 2 is rejected. What to do for a masked edit on the older row.
- MCP outputSchema vs Sume output_schema: who sets the contract
In MCP the server declares a tool's outputSchema. In a Sume Agent Completion you send output_schema per run, and the result can still come back degraded.
- MCP tool-name rules (2025-11-25): do Sume's tool ids comply?
The MCP 2025-11-25 spec says tool names should be 1-128 characters from A-Z, a-z, 0-9, underscore, hyphen and dot. Sume's longest documented id is 36.
Written by Sume