Decagon says Chord is 90% indistinguishable: run your own blind test

Decagon cites about 90% indistinguishability across three voice samples. How to run a fair blind listening test on any TTS output, with a cost estimate.

5 min readSume
All posts

Decagon's Voice 3 announcement says Chord reported that in a blind test across three voices, where listeners picked the human in each pair, roughly 90% on average could not tell. Treat it as a vendor claim from a small sample, and run your own test on your own script, which takes about an hour and a few cents.

What the claim does and does not tell you

The page names three voice samples. It does not tell a buyer your language, your script, or how the listeners were chosen. A claim about customer-conversation voices also says little about a 30-second product ad read, which has different pacing and emotion.

So use the number as a reason to listen, not as a ranking.

A fair blind test in five steps

Keep it small and honest.

  • Pick one script of 3 to 5 sentences you actually ship.
  • Record a human read of it, or use a real human clip you already have.
  • Generate the same script with each candidate voice. In Sume that is one tts_create job per voice.
  • Rename every file to a random code and randomize the order.
  • Ask 10 or more listeners to mark each clip human or synthetic, then count.

What a run costs on Sume

A 400-character script is about 27 seconds at 15 characters per second, which is a planning assumption, not a Sume figure. At the rate card's $0.0475 per 1,000 characters it rounds up to 2 cents per job. Four voices cost about 8 cents. Confirm the live figure in GET /v1/catalog.

Use dry_run and max_spend_usd on the MCP tool to preview the cost and cap it before the batch runs.

Example blind test budget (read 2026-10-07)
ItemCountCost
Candidate voices44 jobs
Script length400 characters2 cents each at $0.0475 per 1,000
Total TTS spend4 jobsabout 8 cents
Listeners10 or moreNo cost

Reading the result

If listeners cannot tell a clip from the human read at a rate near chance, that voice is good enough for that script. If they can, listen for what gave it away: breath, pacing on numbers, or flat endings. Then change the script, not just the voice.

Run the test again with the real final audio before launch. Two different scripts can swap the winner.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume