Voice agent layers: which Oct 2026 launch fills STT, LLM, TTS

A voice agent has three layers. MAI-Transcribe-2-Streaming, Mercury Voice and MAI-Voice-2.1 or Pocket TTS fill them, and Sume covers the non-live work.

6 min readSume
All posts

Which of this week's launches is which part of a voice agent? A voice agent has three live layers: speech to text, a language model, and text to speech. MAI-Transcribe-2-Streaming is the first, Inception's Mercury Voice is the second, and MAI-Voice-2.1 or Kyutai Pocket TTS is the third.

No single launch covers all three, and each vendor states its own latency on its own basis, so the figures below are not meant to be added up.

One launch per layer

Each row is a vendor claim from the vendor's own page. Treat the latency figures as starting points for your own test, since they are measured on different things: first partial text, first answer token, end-to-end synthesis, and first audio chunk.

Voice-agent layers and the launches that fill them (read 2026-10-03)
LayerLaunchStated figureAccess
Speech to textMAI-Transcribe-2-StreamingPartials in just over 100 ms; 60 languagesAvailable, $0.54 per hour introductory through year end
Language modelMercury Voice320 ms median and 750 ms p95 to first answer tokenGenerally available for enterprise customers from October 1, 2026
Text to speechMAI-Voice-2.1-Flash150 ms end to end for 45 s of audio, per MicrosoftPublic preview, no SLA
Text to speech, localKyutai Pocket TTSAbout 200 ms to first chunk; about 6x real time on an M4 CPUOpen, MIT license

What each source says it covers

Microsoft's announcement gives the transcriber figure and the Flash figure. Its Learn page places Flash as best for real-time voice agents, call-center and IVR flows, and says the full MAI-Voice-2.1 prioritizes naturalness over latency-critical scenarios.

The Inception announcement says Mercury Voice is the language-model step, reached through an OpenAI-compatible endpoint, and that it is generally available for enterprise customers. The Pocket TTS README describes a 100M-parameter model that runs on CPU.

What none of them cover

Turn-taking, barge-in, telephony transport and tool calls live in your own code or an orchestration platform.

The three layers also say nothing about work that is not live: narrating a recorded video, scoring it, captioning it, or transcribing a finished call. That is batch work, and a job interface fits it better than a stream.

Where Sume fits

Sume's audio endpoints are jobs. You submit text or a file, and read the result by polling or webhook; synchronous mode waits at most 30 seconds, as set out in Sume jobs and results. That makes them a fit for narration, music, captions and post-call transcripts, and a poor fit for a live conversation, where each turn must stream.

A practical split: use a streaming stack for the call itself, then send the recording to a job-based path for the clean transcript, the captioned clip or the narrated recap that you actually publish.

How to test the stack

Measure each layer on your own audio and prompts, from the point your system receives input to the point it has output, and record the median and the 95th percentile separately. A stack is only as fast as its slowest common path, and the tail matters more on a phone call than the average. The latency budget post sets out one way to do it.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume