AI voice agent latency budget: who owns which 100 milliseconds

Vendor numbers for speech-to-text, the language model and text-to-speech side by side, with a note on where an async file API like Sume belongs and where not.

5 min readSume
All posts

A voice agent hears speech, thinks, and speaks, and the three stages each have a vendor number this week. Microsoft says MAI-Transcribe-2-Streaming produces its first partials in just over 100 ms. Inception says Mercury Voice answers with a median time to first token under 320 ms and 750 ms at p95. Microsoft lists MAI-Voice-2.1-Flash at 150 ms end to end for 45 seconds of audio, and Kyutai's README gives about 200 ms to the first Pocket TTS chunk.

Add them naively and you get a figure under a second, but these are vendor numbers measured under different conditions, so treat them as a way to see where the budget goes, not as a prediction. Sume belongs in this picture only on the edges: it is an async job API, not a live stage.

What does each vendor claim?

Each row is that vendor's own page, read 2026-10-03. The clocks start and stop at different points, which is why the last column matters.

Latency claims from vendor pages, read 2026-10-03: Microsoft AI, Inception, Kyutai README.
StageProductClaimed numberWhat the clock measures
Speech to textMAI-Transcribe-2-StreamingFirst partials in just over 100 msAudio in to first provisional words out
Language modelMercury VoiceMedian under 320 ms, p95 750 msTime to first answer token on production voice prompts, per Inception
Text to speechMAI-Voice-2.1-Flash150 ms end to end for 45 s of audioMicrosoft's wording; not defined further on the page
Text to speechPocket TTS (self-hosted)About 200 ms to the first chunkOn the README's test hardware

Where does the budget really go?

The language model is the biggest single number in that table, which is why Inception built a voice-specific one. The other two stages are already near 100-200 ms in the vendors' own figures. The bigger risk is the sum of everything else: network hops between three vendors, the end-of-speech decision (Microsoft's streaming model has no server-side turn detection, so your code decides when to commit), and the time before the first audio byte reaches a phone line.

  • Decide who ends the turn. With MAI-Transcribe-2-Streaming the client sends commit, so your voice-activity logic is part of the latency.
  • Start text to speech on the first sentence of the model's answer, not the last.
  • Measure time to first audio at the caller, not at the vendor's edge.
  • Keep fixed phrases off the live path entirely; see the pre-rendered prompts post linked below.

Where does Sume fit, and where does it not?

Sume STT and TTS are jobs. TTS 1.0 is explicitly non-streaming: a finished audio file arrives after the whole clip is synthesised, and mode: sync waits up to 30 seconds at most. That is far outside a sub-second turn, so do not put it between a caller's sentence and the agent's reply.

It does fit around the call. Render the greeting, hold message and fixed confirmations ahead of time as files (see Pre-render voice agent greetings), then transcribe the recording afterwards for QA with word timings. For a deeper look at the TTS side, read Low latency text to speech: 150 ms streaming vs async jobs.

What can you move off the live path?

Every millisecond you remove from the live path comes from doing something earlier. Fixed lines can be rendered ahead of time. The caller's number, language and account tier can be looked up before the first question, so the first model call does not wait for a database. Tool results that rarely change can be cached for the length of a call.

Another saving is deciding what the agent says first. Many agents open with a short filler ("Sure, one moment") while a slower answer is prepared; that filler can be a pre-rendered file played the moment the caller stops talking, with the real answer streamed after it. The caller hears a response in the time it takes to start a file, and the model gets a few hundred more milliseconds to finish.

None of this changes the vendor numbers in the table, but it changes what the caller experiences, which is the number that matters. Time it end to end with your own recordings, as above, after each change.

How do you check your own budget?

Record ten real calls, timestamp three moments in each (caller stops speaking, agent's first audio byte, agent's first full word), and compute the median and the slowest. A vendor number that looks great alone often gains 200-400 ms in your network path, and only your own recording shows it. Run the same timing again after you move any one stage, so you learn which vendor swap actually helped.

Sources

Related posts

More in Models

All Models posts

Written by Sume