MAI-Voice 50.3% of 4,000 listeners: how to quote the Turing claim

Microsoft said 50.3% of 4,000 listeners rated MAI-Voice as equally or more human-like than human recordings. What it covers, and a cheap test.

5 min readSume
All posts

The sentence, piece by piece

Microsoft's claim is narrow: in a 4,000-listener Turing test, 50.3 percent of listeners rated MAI-Voice as equally or more human-like than human recordings. If you write about it, quote it that way, because each piece of the sentence limits what it proves.

The source for the number is the Unite.AI report on the October 1, 2026 launch (read 2026-10-07), which attributes the result to Microsoft. This post does not have the original study, so everything below is about how to read the sentence, not about the study design.

What the number does not say

  • It is Microsoft's statement, relayed by a third-party report. Say "Microsoft said".
  • The test combined the two new voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The report gives no separate figure for each model, so do not assign the number to one of them.
  • "Equally or more human-like" merges two outcomes: listeners who preferred the synthetic voice and listeners who could not tell. The report does not split them.
  • Half of 4,000 is about 2,000 people. As plain arithmetic, a score within a point or two of 50 percent is close to a coin flip, so the claim reads as "hard to tell apart", not as "better than people".

Test it on your own script instead

A listener test for the voice you plan to use takes an afternoon and costs less than a sandwich. Generate the same ten lines in two or three candidate voices, put the files in a blind A/B sheet, and ask five colleagues which one sounds like a person reading.

On Sume, text-to-speech is metered per character at the public rate, rounded up to a whole cent per job. Ten lines of 300 characters each are 3,000 characters. Each job is 300 characters at $47.50 per million, or 1.425 cents, which rounds to 2 cents. Ten lines in one voice therefore cost 20 cents, and 60 cents for three voices. Spaces and punctuation count toward the 300.

The table sets the per-line cost beside the two list prices from the launch report, so you can see how the cent rounding changes the picture for very short lines.

Cost of ten 300-character test lines, one voice (list prices read 2026-10-07)
OptionPrice basisTen lines (3,000 characters)
MAI-Voice-2.1$22 per 1M characters, per Unite.AI$0.066 before any rounding
MAI-Voice-2.1-Flash$15 per 1M characters, per Unite.AI$0.045 before any rounding
Sume TTS$47.50 per 1M characters, 1 cent minimum step per job$0.20 (ten jobs of 2 cents)

Reading the cost row

The MAI figures are straight multiplication of the quoted list price; the report does not describe a minimum charge, so none is assumed. The Sume row follows the per-job rounding.

One more point for any public claim: if you ship synthetic voice to listeners, say so in the place they will see it. A Turing result describes how a listening panel scored a sample. It is not a reason to leave out a disclosure line.

Sources

Related posts

More in Models

All Models posts

Written by Sume