MAI-Voice 50.3% of 4,000 listeners: how to quote the Turing claim
Microsoft said 50.3% of 4,000 listeners rated MAI-Voice as equally or more human-like than human recordings. What it covers, and a cheap test.

The sentence, piece by piece
Microsoft's claim is narrow: in a 4,000-listener Turing test, 50.3 percent of listeners rated MAI-Voice as equally or more human-like than human recordings. If you write about it, quote it that way, because each piece of the sentence limits what it proves.
The source for the number is the Unite.AI report on the October 1, 2026 launch (read 2026-10-07), which attributes the result to Microsoft. This post does not have the original study, so everything below is about how to read the sentence, not about the study design.
What the number does not say
- It is Microsoft's statement, relayed by a third-party report. Say "Microsoft said".
- The test combined the two new voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The report gives no separate figure for each model, so do not assign the number to one of them.
- "Equally or more human-like" merges two outcomes: listeners who preferred the synthetic voice and listeners who could not tell. The report does not split them.
- Half of 4,000 is about 2,000 people. As plain arithmetic, a score within a point or two of 50 percent is close to a coin flip, so the claim reads as "hard to tell apart", not as "better than people".
Test it on your own script instead
A listener test for the voice you plan to use takes an afternoon and costs less than a sandwich. Generate the same ten lines in two or three candidate voices, put the files in a blind A/B sheet, and ask five colleagues which one sounds like a person reading.
On Sume, text-to-speech is metered per character at the public rate, rounded up to a whole cent per job. Ten lines of 300 characters each are 3,000 characters. Each job is 300 characters at $47.50 per million, or 1.425 cents, which rounds to 2 cents. Ten lines in one voice therefore cost 20 cents, and 60 cents for three voices. Spaces and punctuation count toward the 300.
The table sets the per-line cost beside the two list prices from the launch report, so you can see how the cent rounding changes the picture for very short lines.
| Option | Price basis | Ten lines (3,000 characters) |
|---|---|---|
| MAI-Voice-2.1 | $22 per 1M characters, per Unite.AI | $0.066 before any rounding |
| MAI-Voice-2.1-Flash | $15 per 1M characters, per Unite.AI | $0.045 before any rounding |
| Sume TTS | $47.50 per 1M characters, 1 cent minimum step per job | $0.20 (ten jobs of 2 cents) |
Reading the cost row
The MAI figures are straight multiplication of the quoted list price; the report does not describe a minimum charge, so none is assumed. The Sume row follows the per-job rounding.
One more point for any public claim: if you ship synthetic voice to listeners, say so in the place they will see it. A Turing result describes how a listening panel scored a sample. It is not a reason to leave out a disclosure line.
Sources
Related posts
More in Models
- Three characters, three dances: Omni IMAGE_REF and VIDEO_REF tokens
Google's Omni 1.1 demo swaps three dancers for a dog, an octopus and a bear. Here is the same request on Sume, with the 0-based reference tokens in order.
- MiniMax H3 or H3 Max after Sora: native 768p vs latent 1080p prices
On Sume, H3 renders native 480p or 768p and bills 2K/4K upscales; H3 Max adds 1080p as a latent refinement of 768p. 10-second prices side by side.
- Minimum clip length by Sume video model: 2, 3, 4 or 5 seconds
Wan 3.0 starts at 2 s, Gemini Omni Flash 1.1 at 3 s, Seedance and Genjutsu at 4 s, MiniMax H3 and Recast at 5 s. Full min and max table for a Sora port.
- Music Router image_url: public HTTPS only, and null to clear it
Sume music image_url must be a public HTTPS image, and null clears it on a reused request object. When a mood still helps and when the prompt does more.
Written by Sume