Eleven v4 says 75% prefer it: run your own 20-listener test

ElevenLabs reports ~75% preference for Eleven v4 in blind tests. A pre-named voice winning 15 of 20 is ~2% by chance. Run your own test on Sume.

4 min readSume
All posts

ElevenLabs says Eleven v4 was ranked first by Artificial Analysis in September 2026 and was preferred by about 75% of listeners in blind head-to-head tests (read 2026-10-05). Those are the vendor's tests, with its own scripts and listeners. Your audience and your script may not behave the same way, and a voice for your ad needs your answer, not an average.

How many listeners is enough

If voices were equal, each listener would pick either at random. With 20 listeners, 15 or more picking one voice you named in advance happens by chance about 2% of the time (21,700 of 1,048,576 possible outcomes, a one-sided sign test). If either voice could be the winner, the chance doubles to about 4%. 14 of 20 is not enough to call. If your audience is small, a 20-person test can only prove a big gap.

Set up the test

  • Pick 5 lines from the real script, 300 to 500 characters each.
  • Render each line with voice A and voice B, using the same language setting.
  • Name files with random codes and keep the key in a separate sheet.
  • Present pairs in random order, with A first half the time.
  • Ask one question: which would you rather hear in the ad?

What it costs on Sume

10 jobs of 400 characters each is 4,000 characters, or $0.19 at $0.0475 per 1,000 characters. Each job rounds to a cent: $0.02 each, $0.20 total. Sume's catalog today is the Sonic family through the TTS Router, so the test compares Sume voices with each other.

Read the results honestly

Count wins per line, not only in total. A voice that wins four lines and loses one on a brand name tells you about pronunciation, not taste. Keep the files and the sheet; if you change voices later, you have a baseline.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume