Voice agent TTS latency: Eleven v4 Turbo 150ms vs Cartesia Sonic

Eleven v4 Turbo claims about 150ms to first speech; Cartesia Sonic-3.6 claims under 90ms. Not like-for-like; Sume TTS is not for live agents.

4 min readSume
All posts

On the vendors' own numbers, Cartesia Sonic-3.6 (under 90ms to first audio) is ahead of Eleven v4 Turbo (about 150ms median time to first speech). Do not rank them on that gap alone. The two pages define and measure the figure differently, and neither states the network path or test region. Measure from your own region, with your own text lengths, and compare p50 and p95.

One more point if you are on Sume: the Sume TTS Router is an asynchronous job API with no streaming TTS in v1, so it suits produced audio such as ads, narration and clips, not live conversation.

What each vendor claims

Vendor pages, read 2026-10-05
ItemEleven v4 TurboCartesia Sonic-3.6
Headline latencyMedian time to first speech about 150msReplies under 90ms to first audio
Other latency figureMedian inference latency about 100msGenerates 132 characters per second vs 68 for v3 Conversational (vendor claim)
StreamingTurbo bidirectional streamingNot covered in this detail on the page read
Launch dateSept 28, 2026Aug 27, 2026
Quality claims#1 on Artificial Analysis, Sept 2026 (vendor claim)#1 on Artificial Analysis for two voice categories (vendor claim)

Why the numbers do not line up

  • Time to first speech, inference latency and time to first audio are three different stopwatches. A median is also not a tail: voice agents feel the p95.
  • A model number excludes your telephony or browser hop, your LLM turn and your endpointing. Those usually dominate the total.
  • Vendor figures are measured on vendor infrastructure. Yours will differ by region and load.

Where Sume fits

Per the Sume TTS Router reference, v1 has no streaming TTS and no Eleven or OpenAI engines. You submit a job, poll /v1/jobs/:id/status and read /v1/jobs/:id/result; the output is a finished audio file. That is the right shape for voiceover on a timeline, and the wrong shape for a phone agent that must start speaking while the model is still writing the sentence.

The Sume docs give no latency figure for TTS jobs, so this post gives none.

A fair test

  • Fix 20 utterances from your real agent: 5 short acknowledgements, 10 sentences, 5 long answers.
  • Log request start, first audio byte and last byte from the same machine for each vendor.
  • Report p50 and p95 per length bucket, and repeat at two times of day.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume