Voice agent TTS latency: Eleven v4 Turbo 150ms vs Cartesia Sonic
Eleven v4 Turbo claims about 150ms to first speech; Cartesia Sonic-3.6 claims under 90ms. Not like-for-like; Sume TTS is not for live agents.

On the vendors' own numbers, Cartesia Sonic-3.6 (under 90ms to first audio) is ahead of Eleven v4 Turbo (about 150ms median time to first speech). Do not rank them on that gap alone. The two pages define and measure the figure differently, and neither states the network path or test region. Measure from your own region, with your own text lengths, and compare p50 and p95.
One more point if you are on Sume: the Sume TTS Router is an asynchronous job API with no streaming TTS in v1, so it suits produced audio such as ads, narration and clips, not live conversation.
What each vendor claims
| Item | Eleven v4 Turbo | Cartesia Sonic-3.6 |
|---|---|---|
| Headline latency | Median time to first speech about 150ms | Replies under 90ms to first audio |
| Other latency figure | Median inference latency about 100ms | Generates 132 characters per second vs 68 for v3 Conversational (vendor claim) |
| Streaming | Turbo bidirectional streaming | Not covered in this detail on the page read |
| Launch date | Sept 28, 2026 | Aug 27, 2026 |
| Quality claims | #1 on Artificial Analysis, Sept 2026 (vendor claim) | #1 on Artificial Analysis for two voice categories (vendor claim) |
Why the numbers do not line up
- Time to first speech, inference latency and time to first audio are three different stopwatches. A median is also not a tail: voice agents feel the p95.
- A model number excludes your telephony or browser hop, your LLM turn and your endpointing. Those usually dominate the total.
- Vendor figures are measured on vendor infrastructure. Yours will differ by region and load.
Where Sume fits
Per the Sume TTS Router reference, v1 has no streaming TTS and no Eleven or OpenAI engines. You submit a job, poll /v1/jobs/:id/status and read /v1/jobs/:id/result; the output is a finished audio file. That is the right shape for voiceover on a timeline, and the wrong shape for a phone agent that must start speaking while the model is still writing the sentence.
The Sume docs give no latency figure for TTS jobs, so this post gives none.
A fair test
- Fix 20 utterances from your real agent: 5 short acknowledgements, 10 sentences, 5 long answers.
- Log request start, first audio byte and last byte from the same machine for each vendor.
- Report p50 and p95 per length bucket, and repeat at two times of day.
Sources
Related posts
More in Comparisons
- ElevenLabs Ads Engine localizes ads in 50+ languages; what Sume covers
Ads Engine pulls ads from Google, Meta and LinkedIn and localizes them. Sume has no ad-account link, but burns translated captions at $0.20 a language.
- ElevenLabs Lip Sync at 480p or 720p vs Sume's lip-sync routes
ElevenLabs' lip-sync tool takes a chest-up portrait plus speech, with 480p or 720p output. Sume has Fabric and H3 Max lip-sync by the second. How they differ.
- ElevenLabs Music 'Commercial, no streaming': what an ad can use
ElevenLabs lists Starter as 'Commercial, no streaming' and Creator as 'Commercial, not enterprise'. What its terms exclude: film, TV, radio, and Studio Games.
- ElevenLabs-UMG vs Suno-Warner: licensed AI music models compared
ElevenLabs signed a multi-year deal with UMG on Sept 10; Suno shipped v6 on licensed data a day earlier. What each side has said and what is still unreleased.
Written by Sume