Speech Arena Elo 1,319: what the score means
Eleven v4 reportedly leads Artificial Analysis' Provider Voice Arena at Elo 1,319. What an arena Elo tells you and how to run your own blind test.

An OrcaRouter article dated Sep 28, 2026 puts Eleven v4 first on Artificial Analysis' Provider Voice Arena at an Elo of 1,319, ahead of Cartesia Sonic 3.6 at 1,276. An arena Elo ranks voices by how listeners voted in pairwise comparisons; it does not tell you which voice suits your script.
The reported figures
These numbers come from the OrcaRouter article, read on the date below. They describe one leaderboard on one day.
| Item | Reported value |
|---|---|
| Eleven v4 on the Provider Voice Arena | First, Elo 1,319 |
| Cartesia Sonic 3.6 (next named model) | Elo 1,276 |
| Eleven v4 Turbo latency | About 150 ms median to first speech (the article notes latency is not comparable across vendors) |
How to read an arena Elo
An Elo score is relative. It comes from head-to-head votes, so a lead means listeners preferred that voice more often than the others in the test prompts used. The gap between two scores matters more than the absolute figure, and a small gap can sit within noise.
Arena prompts are general. A leaderboard cannot know your language, accent, pacing, product names or the emotion your script needs. A voice ranked first overall can still mispronounce a brand name.
- Elo is relative, not an absolute quality score.
- Small gaps between neighbours may be noise.
- Prompts are generic, yours are not.
- Latency matters for live use, less for rendered video.
Run your own blind test
Sume exposes tts_create among its paid MCP tools. Write five or six lines from your real script, including names and numbers, and synthesize the same lines with each voice you are considering. Rename the files so you cannot tell which is which, and have two or three colleagues rank them.
Keep the script, the settings and the order of lines identical across voices. Record the winner per line, not only overall; a voice that wins on narration may lose on a short call to action.
blind-test checklist
1. Same 5-6 lines for every voice, with brand names and numbers.
2. Rename outputs to A, B, C before anyone listens.
3. Each listener ranks per line, then overall.
4. Re-run the winner on a full-length script before committing.Sources
Related posts
More in Comparisons
- Spotify AI covers and remixes: licensed tool vs an original score
Spotify and UMG announced a paid add-on for fan covers and remixes. A licensed remix tool is not an original video score. Here is the Sume route for the second.
- Stable Diffusion alternatives in 2026: open weights or hosted API
SD 3.5 is still Stability's newest flagship image model. FLUX.2 [klein], Qwen-Image 2.0 and Z-Image Turbo are the open options. How to pick between them.
- Suno Speech beta: voice and music in one pass, or separate tracks?
Suno's Speech beta makes voice and music in one track. Its blog lists wandering accents and long pauses. When to prefer separate TTS, music and a timeline mix.
- Suno v6 and label partners: what it means for an ad soundtrack
Suno's v6 post names WMG, BMG and Believe and upload screening. What a rights-minded team should check before using any AI music in a paid ad.
Written by Sume