Voice agent layers: which Oct 2026 launch fills STT, LLM, TTS
A voice agent has three layers. MAI-Transcribe-2-Streaming, Mercury Voice and MAI-Voice-2.1 or Pocket TTS fill them, and Sume covers the non-live work.

Which of this week's launches is which part of a voice agent? A voice agent has three live layers: speech to text, a language model, and text to speech. MAI-Transcribe-2-Streaming is the first, Inception's Mercury Voice is the second, and MAI-Voice-2.1 or Kyutai Pocket TTS is the third.
No single launch covers all three, and each vendor states its own latency on its own basis, so the figures below are not meant to be added up.
One launch per layer
Each row is a vendor claim from the vendor's own page. Treat the latency figures as starting points for your own test, since they are measured on different things: first partial text, first answer token, end-to-end synthesis, and first audio chunk.
| Layer | Launch | Stated figure | Access |
|---|---|---|---|
| Speech to text | MAI-Transcribe-2-Streaming | Partials in just over 100 ms; 60 languages | Available, $0.54 per hour introductory through year end |
| Language model | Mercury Voice | 320 ms median and 750 ms p95 to first answer token | Generally available for enterprise customers from October 1, 2026 |
| Text to speech | MAI-Voice-2.1-Flash | 150 ms end to end for 45 s of audio, per Microsoft | Public preview, no SLA |
| Text to speech, local | Kyutai Pocket TTS | About 200 ms to first chunk; about 6x real time on an M4 CPU | Open, MIT license |
What each source says it covers
Microsoft's announcement gives the transcriber figure and the Flash figure. Its Learn page places Flash as best for real-time voice agents, call-center and IVR flows, and says the full MAI-Voice-2.1 prioritizes naturalness over latency-critical scenarios.
The Inception announcement says Mercury Voice is the language-model step, reached through an OpenAI-compatible endpoint, and that it is generally available for enterprise customers. The Pocket TTS README describes a 100M-parameter model that runs on CPU.
What none of them cover
Turn-taking, barge-in, telephony transport and tool calls live in your own code or an orchestration platform.
The three layers also say nothing about work that is not live: narrating a recorded video, scoring it, captioning it, or transcribing a finished call. That is batch work, and a job interface fits it better than a stream.
Where Sume fits
Sume's audio endpoints are jobs. You submit text or a file, and read the result by polling or webhook; synchronous mode waits at most 30 seconds, as set out in Sume jobs and results. That makes them a fit for narration, music, captions and post-call transcripts, and a poor fit for a live conversation, where each turn must stream.
A practical split: use a streaming stack for the call itself, then send the recording to a job-based path for the clean transcript, the captioned clip or the narrated recap that you actually publish.
How to test the stack
Measure each layer on your own audio and prompts, from the point your system receives input to the point it has output, and record the median and the 95th percentile separately. A stack is only as fast as its slowest common path, and the tail matters more on a phone call than the average. The latency budget post sets out one way to do it.
Sources
Related posts
More in Comparisons
- Voice isolator vs audio detach: 500 MB and 1 hour vs 1800 seconds
ElevenLabs Voice Isolator removes background noise from speech. Sume audio detach extracts the video audio track as is; it does not separate vocals from music.
- Wan 3.0 720p 5s clip: Pika 33 credits, Runway API 50, Sume $0.625
Pika lists Wan 3.0 720p 5s at 33 credits ($0.29-$0.55 by plan or pack); Runway's API charges 50 credits ($0.50). Sume bills $0.625 at list x 1.25.
- Wan 3.0 or MiniMax H3 Max: which to pin for reference-to-video
Both take image, video and audio references on Sume. Wan runs 2 to 30 seconds with 5 reference videos; H3 Max runs 5 to 15 seconds with always-on stereo audio.
- Wan 3.0 or Seedance 2.5: which is cheaper for a 10-second 720p clip?
Wan 3.0 is cheaper: a 10-second 720p clip is $1.00 at fal list ($1.25 on Sume) against about $4.62 ($5.78 on Sume) for Seedance 2.5, roughly 4.6x.
Written by Sume