Speech Arena Elo 1,319: what the score means

Eleven v4 reportedly leads Artificial Analysis' Provider Voice Arena at Elo 1,319. What an arena Elo tells you and how to run your own blind test.

4 min readSume
All posts

An OrcaRouter article dated Sep 28, 2026 puts Eleven v4 first on Artificial Analysis' Provider Voice Arena at an Elo of 1,319, ahead of Cartesia Sonic 3.6 at 1,276. An arena Elo ranks voices by how listeners voted in pairwise comparisons; it does not tell you which voice suits your script.

The reported figures

These numbers come from the OrcaRouter article, read on the date below. They describe one leaderboard on one day.

Provider Voice Arena coverage, Sep 28, 2026 (read 2026-10-03)
ItemReported value
Eleven v4 on the Provider Voice ArenaFirst, Elo 1,319
Cartesia Sonic 3.6 (next named model)Elo 1,276
Eleven v4 Turbo latencyAbout 150 ms median to first speech (the article notes latency is not comparable across vendors)

How to read an arena Elo

An Elo score is relative. It comes from head-to-head votes, so a lead means listeners preferred that voice more often than the others in the test prompts used. The gap between two scores matters more than the absolute figure, and a small gap can sit within noise.

Arena prompts are general. A leaderboard cannot know your language, accent, pacing, product names or the emotion your script needs. A voice ranked first overall can still mispronounce a brand name.

  • Elo is relative, not an absolute quality score.
  • Small gaps between neighbours may be noise.
  • Prompts are generic, yours are not.
  • Latency matters for live use, less for rendered video.

Run your own blind test

Sume exposes tts_create among its paid MCP tools. Write five or six lines from your real script, including names and numbers, and synthesize the same lines with each voice you are considering. Rename the files so you cannot tell which is which, and have two or three colleagues rank them.

Keep the script, the settings and the order of lines identical across voices. Record the winner per line, not only overall; a voice that wins on narration may lose on a short call to action.

blind-test checklist
1. Same 5-6 lines for every voice, with brand names and numbers.
2. Rename outputs to A, B, C before anyone listens.
3. Each listener ranks per line, then overall.
4. Re-run the winner on a full-length script before committing.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume