Voice agent latency claims: 100 ms, 150 ms, 90 ms and Sume TTS jobs

ElevenLabs and Cartesia quote 90 to 150 ms for live voice. What those figures measure, and why Sume TTS jobs are for produced audio rather than live calls.

4 min readSume
All posts

The vendor numbers are time to first audio or median inference latency for a streaming voice agent, so they do not compare to a Sume TTS job. Sume TTS is an asynchronous job API for produced audio such as voiceovers, with a synchronous wait capped at 30 seconds; for a live phone agent you want a vendor streaming endpoint, not a job.

If you are building a real-time agent, the latency claims matter. If you are producing a video voiceover, they are nearly irrelevant. Here is how to tell which side of the line you are on.

What the figures say

The numbers are close but not identical in meaning, which is why a side-by-side table needs a unit column.

Latency figures from vendor pages (read 2026-10-03)
SourceFigureWhat it describes
ElevenLabs v4 blogAbout 150 ms median time to first speechv4 Turbo
ElevenLabs v4 blogAbout 100 ms median inference latencyv4
ElevenLabs v4 Turbo in ElevenAgents postAbout 100 ms median inference latencyv4 Turbo
ElevenLabs models docAbout 75 msFlash v2.5 and Flash v2
ElevenLabs models docAbout 280 msv3 Conversational
Cartesia Sonic pageUnder 90 msTime to first audio

Inference latency is not conversation latency

A median inference figure excludes the network hop, the speech recognition step and the language model's thinking time. ElevenLabs itself frames v4 Turbo as co-optimized with transcription and turn-taking inside ElevenAgents, which suggests the number is meant for the whole agent loop. Treat each vendor figure as a best case measured by that vendor.

What Sume TTS is built for

The TTS 1.0 and TTS Router contracts create a job. You submit a transcript, choose async, sync, subscribe or webhook mode, and receive audio as a result artifact. In sync mode the wait is capped at 30 seconds, after which you poll or subscribe. Output can be mp3, wav or raw with several sample rates, and the transcript limit is 20,000 characters per request.

That shape suits voiceovers, ad reads and narrated explainers, where you need good audio, word timings and a stored result you can reuse. It does not suit a caller waiting for a reply, because a job has queueing, storage and settlement steps that a streaming socket does not.

  • Live agent: choose a streaming provider endpoint outside Sume.
  • Produced audio: use TTS 1.0 or the Router, and take word timestamps if you will burn captions.
  • Hybrid: generate the agent's greeting and fixed prompts on Sume ahead of time, then stream the dynamic parts elsewhere.

A quick decision rule

Ask whether a human is waiting in real time. If yes, latency per word matters and a job API is the wrong tool. If no, spend your attention on voice quality, language, pronunciation and cost. The voice agent cost post covers the budget side of the live case.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume