MAI-Transcribe-2-Streaming has no server VAD: who ends the turn?

In the preview, turn_detection only accepts null, so your client sends the commit event. Sume STT has no live socket; it cuts sentences after the job finishes.

4 min readSume
All posts

With MAI-Transcribe-2-Streaming, your client decides when a turn ends. The realtime page says turn_detection supports only null, so there is no server-side voice activity detection and no auto-commit; you send input_audio_buffer.commit yourself. Sume STT 1.0 has no live socket at all: you submit an audio file, wait for the job, and get word timings plus optional sentence segments back.

What Microsoft documents for the preview

The realtime page marks the model as public preview with no SLA. It also says noise_reduction supports only null, and that session settings cannot change after the first audio append. The practical consequence is that silence handling, pause detection and turn boundaries are code you write on your side of the socket.

The announcement pitches first partials in just over 100 ms and $0.54 per hour through the end of the year. Neither number tells you where a turn ends.

Turn handling, MAI streaming preview vs Sume STT 1.0 (Microsoft pages read 2026-10-05; Sume from the repo)
QuestionMAI-Transcribe-2-StreamingSume STT 1.0
Who ends a turnYour client, via input_audio_buffer.commitNobody live; sentence segments are cut after the job
Server VADturn_detection accepts only nullNot applicable; no live audio
Settings after first audioLocked for the sessionFixed per request
TransportWebSocket, one session of up to one hourPOST /v1/stt-1.0/transcribe, then poll the job
Partial textdelta and intermediate eventsNone; the result arrives complete

Where Sume's equivalent of a turn boundary comes from

If you pass segmentation: {"mode": "sentence"} to Sume STT, the result carries gapless segments[] built from the timed words. A segment ends on sentence-final punctuation; when the transcript has no punctuation at the tail, a silence of at least 0.5 seconds splits it. That is a post-hoc cut, useful for captions, not a substitute for a live endpointer.

The job is asynchronous, so see jobs and results for polling. The segments[] drop into caption lines; the video captions docs cover cues and segments input.

If you build commit logic yourself

A few decisions come with owning the endpointer. They are yours in the preview, so write them down before the first demo.

  • Pick a silence threshold in milliseconds and test it on your noisiest real audio, since noise reduction is also off.
  • Decide whether a commit fires on silence, on a push-to-talk release, or on a client timer.
  • Remember that settings are frozen after the first audio append, so the language choice must be made before you stream.
  • Log every commit with a timestamp, so a late or missing turn can be traced to your client rather than the model.

Which one to pick

Pick MAI streaming if you are building a voice interface where the user is waiting on partial text, and you are willing to write your own commit logic against a preview with no SLA. Pick Sume STT when the audio already exists as a file and you want timed words, sentence segments or captions out of it.

When not to use Sume STT: anything interactive. A job is not a conversation turn, and the request caps audio at 10 minutes (duration_seconds 1 to 600).

Sources

Related posts

More in Models

All Models posts

Written by Sume