TTS leaderboard: 33 Elo points rank 5 to 12
On Versely's September 2026 voice leaderboard, ranks 5 to 12 span 33 Elo points. Here is what that gap means and how to test a voice with Sume's tts_create.

The gap between rank 5 and rank 12 on Versely's September 2026 text-to-speech leaderboard is 33 Elo points: Gemini 3.1 Flash TTS sits at 1212 and ElevenLabs at 1179. A gap that small is a reason to audition voices on your own script rather than pick by rank.
What the leaderboard lists
The Versely write-up gives these positions. The 33-point figure is simple subtraction of the two Elo values.
| Model | Rank | Elo |
|---|---|---|
| Gemini 3.1 Flash TTS | 5 | 1212 |
| Cartesia Sonic 3.5 | 7 | 1203 |
| Inworld TTS 2 | 10 | 1187 |
| ElevenLabs | 12 | 1179 |
Reading a 33-point spread
Versely also notes that Gemini 3.1 Flash TTS shipped on April 15, 2026 with 30 voices and multi-speaker dialogue. Those are feature facts, not quality facts. An Elo number summarizes pairwise listener votes across many prompts. It does not tell you how a model reads your product names, your language mix or your pacing.
Treat the table as a shortlist. Anything within a few dozen points of the top is worth one blind test with 3 to 5 of your real lines.
Testing a voice through Sume
Sume's hosted MCP exposes tts_create as a paid generation tool, next to stt_create and music_create. Paid calls need an idempotency_key, and dry_run=true previews admission and cost without submitting a job. Ask the server for the live contract with tools_schema and name: "tts_create" before you build on it, because the docs say not to assume parity with the HTTP API.
Which voices or engines a session can pick is listed by that schema, not by this post. Run the same three lines through each option you can select, play them back to back without labels, and keep a note of which one your listeners prefer.
A simple audition routine
- Pick lines that contain numbers, brand names and one long sentence.
- Generate each line once per voice and keep the job ids and settings next to the audio.
- Re-run the winner on a second day to check that results are stable.
- Only then compare price per character or per minute.
Sources
Related posts
More in Comparisons
- Udio downloads are off: export-ready music for video work
Udio disabled downloads after its UMG settlement. Where to get an exportable AI music bed for video instead: Sume's Music Router returns a file URL.
- Reading a vendor-run TTS leaderboard: Gemini 3.8 Flash TTS on VoiceEQ
Hume's blog lists Gemini 3.8 Flash TTS atop its Real-World VoiceEQ board. What that does and does not tell you, plus a blind test to run on your script.
- Veo 3.1 Lite at half of Fast vs an Omni Flash draft
Google says Veo 3.1 Lite costs under half of Veo 3.1 Fast. On Sume the Google video id is gemini-omni-flash-1.1, 3-10 s at 360p to 4K. Compare before budgeting.
- Reference limits: Veo 3.1, Gemini Omni Flash and Sume Video 1.0
Veo 3.1 takes up to 3 reference images, Omni Flash up to 3 clips of 3 s each, and Sume Video 1.0 takes 1 to 9 images. A table to pick by what you need to pin.
Written by Sume