MAI-Transcribe-2-Streaming has no server VAD: who ends the turn?
In the preview, turn_detection only accepts null, so your client sends the commit event. Sume STT has no live socket; it cuts sentences after the job finishes.

With MAI-Transcribe-2-Streaming, your client decides when a turn ends. The realtime page says turn_detection supports only null, so there is no server-side voice activity detection and no auto-commit; you send input_audio_buffer.commit yourself. Sume STT 1.0 has no live socket at all: you submit an audio file, wait for the job, and get word timings plus optional sentence segments back.
What Microsoft documents for the preview
The realtime page marks the model as public preview with no SLA. It also says noise_reduction supports only null, and that session settings cannot change after the first audio append. The practical consequence is that silence handling, pause detection and turn boundaries are code you write on your side of the socket.
The announcement pitches first partials in just over 100 ms and $0.54 per hour through the end of the year. Neither number tells you where a turn ends.
| Question | MAI-Transcribe-2-Streaming | Sume STT 1.0 |
|---|---|---|
| Who ends a turn | Your client, via input_audio_buffer.commit | Nobody live; sentence segments are cut after the job |
| Server VAD | turn_detection accepts only null | Not applicable; no live audio |
| Settings after first audio | Locked for the session | Fixed per request |
| Transport | WebSocket, one session of up to one hour | POST /v1/stt-1.0/transcribe, then poll the job |
| Partial text | delta and intermediate events | None; the result arrives complete |
Where Sume's equivalent of a turn boundary comes from
If you pass segmentation: {"mode": "sentence"} to Sume STT, the result carries gapless segments[] built from the timed words. A segment ends on sentence-final punctuation; when the transcript has no punctuation at the tail, a silence of at least 0.5 seconds splits it. That is a post-hoc cut, useful for captions, not a substitute for a live endpointer.
The job is asynchronous, so see jobs and results for polling. The segments[] drop into caption lines; the video captions docs cover cues and segments input.
If you build commit logic yourself
A few decisions come with owning the endpointer. They are yours in the preview, so write them down before the first demo.
- Pick a silence threshold in milliseconds and test it on your noisiest real audio, since noise reduction is also off.
- Decide whether a commit fires on silence, on a push-to-talk release, or on a client timer.
- Remember that settings are frozen after the first audio append, so the language choice must be made before you stream.
- Log every commit with a timestamp, so a late or missing turn can be traced to your client rather than the model.
Which one to pick
Pick MAI streaming if you are building a voice interface where the user is waiting on partial text, and you are willing to write your own commit logic against a preview with no SLA. Pick Sume STT when the audio already exists as a file and you want timed words, sentence segments or captions out of it.
When not to use Sume STT: anything interactive. A job is not a conversation turn, and the request caps audio at 10 minutes (duration_seconds 1 to 600).
Sources
Related posts
More in Models
- MAI-Voice-2.1 has 23 languages, 26 locales, 28 codes: which to quote
Microsoft says 23 languages and 26 locales; OpenRouter lists 28 codes and says 30+. Quote 23 languages, and use Python to turn the 28 codes into 23.
- MAI-Voice-2.1-Flash: 150ms for 45 seconds of audio, for batch TTS
Microsoft says MAI-Voice-2.1-Flash makes 45s of audio at 150ms end-to-end latency, at $15 per 1M characters. What that does and does not tell a batch TTS user.
- MAI-Voice-2.1-Flash at 45 ms: does a rendered avatar need fast TTS?
Microsoft lists MAI-Voice-2.1-Flash at about 45 ms of inference. A rendered avatar clip does not benefit from it. Where the latency shows up in a Sume job.
- MAI-Voice-2.1 Hindi: six voices listed, and a Hindi request on Sume
Microsoft's MAI-Voice-2.1 page lists six hi-IN voices with different style sets. Compare that with a Hindi request on Sume TTS and what the language field does.
Written by Sume