Decagon says Chord is 90% indistinguishable: run your own blind test
Decagon cites about 90% indistinguishability across three voice samples. How to run a fair blind listening test on any TTS output, with a cost estimate.

Decagon's Voice 3 announcement says Chord reported that in a blind test across three voices, where listeners picked the human in each pair, roughly 90% on average could not tell. Treat it as a vendor claim from a small sample, and run your own test on your own script, which takes about an hour and a few cents.
What the claim does and does not tell you
The page names three voice samples. It does not tell a buyer your language, your script, or how the listeners were chosen. A claim about customer-conversation voices also says little about a 30-second product ad read, which has different pacing and emotion.
So use the number as a reason to listen, not as a ranking.
A fair blind test in five steps
Keep it small and honest.
- Pick one script of 3 to 5 sentences you actually ship.
- Record a human read of it, or use a real human clip you already have.
- Generate the same script with each candidate voice. In Sume that is one tts_create job per voice.
- Rename every file to a random code and randomize the order.
- Ask 10 or more listeners to mark each clip human or synthetic, then count.
What a run costs on Sume
A 400-character script is about 27 seconds at 15 characters per second, which is a planning assumption, not a Sume figure. At the rate card's $0.0475 per 1,000 characters it rounds up to 2 cents per job. Four voices cost about 8 cents. Confirm the live figure in GET /v1/catalog.
Use dry_run and max_spend_usd on the MCP tool to preview the cost and cap it before the batch runs.
| Item | Count | Cost |
|---|---|---|
| Candidate voices | 4 | 4 jobs |
| Script length | 400 characters | 2 cents each at $0.0475 per 1,000 |
| Total TTS spend | 4 jobs | about 8 cents |
| Listeners | 10 or more | No cost |
Reading the result
If listeners cannot tell a clip from the human read at a rate near chance, that voice is good enough for that script. If they can, listen for what gave it away: breath, pacing on numbers, or flat endings. Then change the script, not just the voice.
Run the test again with the real final audio before launch. Two different scripts can swap the winner.
Sources
Related posts
More in Comparisons
- Decagon's Oct 1 launches: which of the four matter to a video team?
Voice 3, Personal Agent Gateway, Agent Modules and Duet Apprentice launched together. Only one touches voice; here is what each does and who needs it.
- Decagon Voice 3 handles 70+ languages mid-call. A TTS job takes one
Voice 3 detects language and switches mid-sentence. A Sume TTS job speaks one language per request. How to plan a multilingual script around that.
- Decagon Voice 3 duplex agent or a voiceover job: which do you need?
Voice 3 listens while it speaks. A voiceover job does not. How to choose between a live voice agent and a file-based TTS job for your project.
- Does Sume have a real-time avatar API? No, here is what it has instead
Sume has no live avatar session. It has async avatar jobs: create an avatar, render a talking video, or lip-sync a still to audio. Routes, limits and prices.
Written by Sume