Two TTS vendors both claim number one: how to settle it yourself

ElevenLabs and Cartesia each cite a top Artificial Analysis spot in September 2026. How to read both claims and run a blind test on Sume with Sonic model ids.

4 min readSume
All posts

Both claims can be true because a leaderboard can have more than one board and more than one date. Do not pick a vendor from the badge; run a blind listening test on your own script. On Sume you can run that test across Sonic 3.6, 3.5 and 3 today, while ElevenLabs Eleven v4 is not a model id you can call on Sume.

This post shows how to read each claim, what the vendors actually wrote, and a small protocol for a fair comparison.

What each vendor wrote

In its v4 Turbo in ElevenAgents post, ElevenLabs says v4 Turbo ranks first on Artificial Analysis as of September 2026, with about 100 ms median inference latency. On its Sonic page, Cartesia says Sonic is first on the Artificial Analysis Speech Arena and on its speech-to-text leaderboards. The Sonic 3.6 post gives Elo scores of 1120 for a controlled voice and 1285 for a provider voice.

The two statements are not the same claim. One concerns a model in a particular product context, the other concerns a leaderboard category, and the Cartesia figures distinguish a controlled voice from a provider voice. Neither page gives you an apples-to-apples score for your script.

Leaderboard claims as stated by each vendor (read 2026-10-03)
VendorClaimWhere
ElevenLabsv4 Turbo ranked first, September 2026v4 Turbo in ElevenAgents post
CartesiaSonic first on the Speech Arena and STT leaderboardsSonic product page
CartesiaElo 1120 controlled voice, 1285 provider voiceSonic 3.6 launch post

Why a badge is not a purchase decision

A speech arena votes on short generic samples. Your product reads product names, prices, addresses and a brand voice in a specific language. A model that wins on conversational English can still stumble on SKU codes in Spanish.

  • Check the category and date behind any badge.
  • Check which voice the score used.
  • Check whether the claim covers your language at all.
  • Treat a latency figure as a claim about the vendor's pipeline, not about your end-to-end time.

A blind test you can run on Sume

Write ten lines from your real script, including at least three with numbers and two with brand names. Generate each through the TTS Router with sonic-3.6 and sonic-3.5 using the same voice and language, and keep the job results, which record the model and voice. Strip the filenames, shuffle, and have three colleagues pick the clip they would publish.

If your shortlist includes Eleven v4, run its clips through ElevenLabs separately and fold them into the same shuffle. Sume cannot generate them, so keep the comparison honest by labelling that arm as made outside Sume.

What to record

Record the model id, voice, language, output format and the listener picks. When a new snapshot ships, repeat the same ten lines. A stable test set tells you more than any ranking page, and it takes an afternoon.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume