Grok STT smart_turn end-of-turn vs Sume STT sentence segmentation

smart_turn predicts when a speaker is done in live streams. Sume STT is batch and cuts sentences from word timings. Which one do you need?

5 min readSume
All posts

Smart Turn and Sume's sentence segmentation solve different problems. Smart Turn, in xAI's streaming Speech to Text, decides live whether a speaker has finished a thought so a voice agent knows when to answer. Sume STT is a batch job: you submit a finished recording and get back text with word timings, optionally grouped into sentences. If you are building a live agent, Sume STT is not the endpoint. If you are subtitling recordings, you do not need end-of-turn prediction at all.

What does smart_turn do?

According to xAI's release notes, the streaming Speech to Text API supports Smart Turn end-of-turn detection. When you enable it with the smart_turn query parameter, an ML model predicts whether the speaker has finished their thought at silence boundaries. A smart_turn_timeout parameter sets a maximum silence fallback. The notes say this reduces false endpointing during dictation, number sequences and pauses. All of this was read on 2026-10-04.

End-of-turn versus sentence segmentation, vendor pages read 2026-10-04
AspectxAI smart_turnSume STT segmentation
ModeStreamingBatch job
Decision madeHas the speaker finished?Where does each sentence start and end?
InputLive audioPublic HTTPS audio_url
Fallback controlsmart_turn_timeoutboundary_lead_ms, 0 to 500, default 70

How does Sume STT segment sentences?

In the Sume OpenAPI spec, segmentation is optional and has one mode, sentence. The sentences are derived from the returned word timings, and the request fails closed with a typed error if the provider returns no timed words. boundary_lead_ms pads each boundary and accepts 0 to 500 with a default of 70.

That makes it useful after the fact: caption cues, clip cutting, and reading-speed checks. See checking caption reading speed from STT segments and cutting a voiceover into sentence clips for how the same field is used on the TTS side.

Can a batch job stand in for end-of-turn detection?

No. A Sume STT job returns when the whole file is processed, so it cannot tell your agent when to speak mid-conversation. Jobs are asynchronous by default, with a bounded wait of at most 30 seconds in sync mode, and are priced at $0.01 per audio minute.

Use it for what happens after the call ends: transcripts, summaries and subtitles. The related question of Microsoft's streaming model is covered in MAI-Transcribe-2 streaming vs Sume job URLs.

What should you pick?

Pick xAI's streaming endpoint with smart_turn for a live voice agent that must stop talking over people. Pick Sume STT for recorded audio where you want sentences with timings. Many teams use both: a streaming model for the call and a batch pass for the archive.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume