Voice agent latency claims: 100 ms, 150 ms, 90 ms and Sume TTS jobs
ElevenLabs and Cartesia quote 90 to 150 ms for live voice. What those figures measure, and why Sume TTS jobs are for produced audio rather than live calls.

The vendor numbers are time to first audio or median inference latency for a streaming voice agent, so they do not compare to a Sume TTS job. Sume TTS is an asynchronous job API for produced audio such as voiceovers, with a synchronous wait capped at 30 seconds; for a live phone agent you want a vendor streaming endpoint, not a job.
If you are building a real-time agent, the latency claims matter. If you are producing a video voiceover, they are nearly irrelevant. Here is how to tell which side of the line you are on.
What the figures say
The numbers are close but not identical in meaning, which is why a side-by-side table needs a unit column.
| Source | Figure | What it describes |
|---|---|---|
| ElevenLabs v4 blog | About 150 ms median time to first speech | v4 Turbo |
| ElevenLabs v4 blog | About 100 ms median inference latency | v4 |
| ElevenLabs v4 Turbo in ElevenAgents post | About 100 ms median inference latency | v4 Turbo |
| ElevenLabs models doc | About 75 ms | Flash v2.5 and Flash v2 |
| ElevenLabs models doc | About 280 ms | v3 Conversational |
| Cartesia Sonic page | Under 90 ms | Time to first audio |
Inference latency is not conversation latency
A median inference figure excludes the network hop, the speech recognition step and the language model's thinking time. ElevenLabs itself frames v4 Turbo as co-optimized with transcription and turn-taking inside ElevenAgents, which suggests the number is meant for the whole agent loop. Treat each vendor figure as a best case measured by that vendor.
What Sume TTS is built for
The TTS 1.0 and TTS Router contracts create a job. You submit a transcript, choose async, sync, subscribe or webhook mode, and receive audio as a result artifact. In sync mode the wait is capped at 30 seconds, after which you poll or subscribe. Output can be mp3, wav or raw with several sample rates, and the transcript limit is 20,000 characters per request.
That shape suits voiceovers, ad reads and narrated explainers, where you need good audio, word timings and a stored result you can reuse. It does not suit a caller waiting for a reply, because a job has queueing, storage and settlement steps that a streaming socket does not.
- Live agent: choose a streaming provider endpoint outside Sume.
- Produced audio: use TTS 1.0 or the Router, and take word timestamps if you will burn captions.
- Hybrid: generate the agent's greeting and fixed prompts on Sume ahead of time, then stream the dynamic parts elsewhere.
A quick decision rule
Ask whether a human is waiting in real time. If yes, latency per word matters and a job API is the wrong tool. If no, spend your attention on voice quality, language, pronunciation and cost. The voice agent cost post covers the budget side of the live case.
Sources
Related posts
More in Comparisons
- Voice agent platforms at 10,000 minutes: eight price lists compared
Telnyx about $596, Deepgram and AssemblyAI $750, Bland $1,400 to $1,499, Retell $700 to $3,100: eight published rates at 10,000 minutes.
- Voice isolator vs audio detach: 500 MB and 1 hour vs 1800 seconds
ElevenLabs Voice Isolator removes background noise from speech. Sume audio detach extracts the video audio track as is; it does not separate vocals from music.
- Wan 3.0 or MiniMax H3 Max: which to pin for reference-to-video
Both take image, video and audio references on Sume. Wan runs 2 to 30 seconds with 5 reference videos; H3 Max runs 5 to 15 seconds with always-on stereo audio.
- Wan 3.0 or Seedance 2.5: which is cheaper for a 10-second 720p clip?
Wan 3.0 is cheaper: a 10-second 720p clip is $1.00 at fal list ($1.25 on Sume) against about $4.62 ($5.78 on Sume) for Seedance 2.5, roughly 4.6x.
Written by Sume