Two TTS vendors both claim number one: how to settle it yourself
ElevenLabs and Cartesia each cite a top Artificial Analysis spot in September 2026. How to read both claims and run a blind test on Sume with Sonic model ids.

Both claims can be true because a leaderboard can have more than one board and more than one date. Do not pick a vendor from the badge; run a blind listening test on your own script. On Sume you can run that test across Sonic 3.6, 3.5 and 3 today, while ElevenLabs Eleven v4 is not a model id you can call on Sume.
This post shows how to read each claim, what the vendors actually wrote, and a small protocol for a fair comparison.
What each vendor wrote
In its v4 Turbo in ElevenAgents post, ElevenLabs says v4 Turbo ranks first on Artificial Analysis as of September 2026, with about 100 ms median inference latency. On its Sonic page, Cartesia says Sonic is first on the Artificial Analysis Speech Arena and on its speech-to-text leaderboards. The Sonic 3.6 post gives Elo scores of 1120 for a controlled voice and 1285 for a provider voice.
The two statements are not the same claim. One concerns a model in a particular product context, the other concerns a leaderboard category, and the Cartesia figures distinguish a controlled voice from a provider voice. Neither page gives you an apples-to-apples score for your script.
| Vendor | Claim | Where |
|---|---|---|
| ElevenLabs | v4 Turbo ranked first, September 2026 | v4 Turbo in ElevenAgents post |
| Cartesia | Sonic first on the Speech Arena and STT leaderboards | Sonic product page |
| Cartesia | Elo 1120 controlled voice, 1285 provider voice | Sonic 3.6 launch post |
Why a badge is not a purchase decision
A speech arena votes on short generic samples. Your product reads product names, prices, addresses and a brand voice in a specific language. A model that wins on conversational English can still stumble on SKU codes in Spanish.
- Check the category and date behind any badge.
- Check which voice the score used.
- Check whether the claim covers your language at all.
- Treat a latency figure as a claim about the vendor's pipeline, not about your end-to-end time.
A blind test you can run on Sume
Write ten lines from your real script, including at least three with numbers and two with brand names. Generate each through the TTS Router with sonic-3.6 and sonic-3.5 using the same voice and language, and keep the job results, which record the model and voice. Strip the filenames, shuffle, and have three colleagues pick the clip they would publish.
If your shortlist includes Eleven v4, run its clips through ElevenLabs separately and fold them into the same shuffle. Sume cannot generate them, so keep the comparison honest by labelling that arm as made outside Sume.
What to record
Record the model id, voice, language, output format and the listener picks. When a new snapshot ships, repeat the same ten lines. A stable test set tells you more than any ranking page, and it takes an afternoon.
Sources
Related posts
More in Comparisons
- Veo 3.1 Lite vs Wan 3.0: price per second at 480p, 720p and 1080p
Google lists Veo 3.1 Lite at $0.05 (720p), $0.08 (1080p) per second; fal lists Wan 3.0 at $0.05 (480p), $0.10 (720p), $0.20 (1080p).
- Veo 3.1 clips are 4, 6 or 8 s: when 10 s needs Omni instead
Veo 3.1 clips are 4, 6 or 8 seconds. For a 10 second shot, Gemini Omni Flash 1.1 on Sume outputs 3 to 10 seconds. A duration table and what stays unavailable.
- Veo 3.1 tiers vs Seedance 2.5: price per second of video
Google lists Veo 3.1 at $0.40 (Standard), $0.10-$0.12 (Fast) and $0.05-$0.08 (Lite) per second; Seedance 2.5 on fal is about $0.46 at 720p. Sume x 1.25: $0.58.
- Veo 3.1 vs Gemini Omni Flash 1.1: what Sume can call
Google lists Veo 3.1 at 1080p and 4K with native audio. On Sume, Gemini Omni Flash 1.1 is the callable Google id: 3-10 s, 360p-4K, audio always on.
Written by Sume