Where to try MAI-Voice-2.1 before paying: a 10-minute listening test
Microsoft lists the MAI Playground, Copilot Audio Expressions and Foundry for MAI-Voice-2.1. Run a ten-minute listening test with a fixed script and cost it.

The fastest free-to-low-cost way to hear MAI-Voice-2.1 is Microsoft's own surfaces. The model page says it is available in the MAI Playground, in Copilot Audio Expressions and in Microsoft Foundry (Azure Speech), and the launch post adds Vercel and OpenRouter (read 2026-10-03). Use the same fixed test script on every engine you consider, because a model that sounds great reading its own demo text can stumble on your product name and your prices.
This post is the script and the scoring, not a review. Listen yourself, and keep the sheet.
Access routes
The model page lists about 550 ms model inference for MAI-Voice-2.1 and about 45 ms for Flash. For a listening test you do not need latency; you need a script that exposes the weak spots.
| Route | Best for | Notes |
|---|---|---|
| MAI Playground | Quick listening | No integration needed |
| Copilot Audio Expressions | Seeing the emotion controls | Named on the model page |
| Microsoft Foundry | Production use inside Azure | $22 per 1M (2.1), $15 per 1M (Flash) |
| Vercel, OpenRouter | Calling it from an existing gateway | Check the gateway's own rate |
A fixed ten-minute test script
Use five lines, each written to break something. A line with numbers and units: a price, a phone number, a date. A line with your brand and product names. A long sentence with two clauses and a list. A short, emotional line the voice should act. And a line in a second language if you need one. Record the same five lines on each engine in the same order.
Score each clip 1 to 5 on pronunciation, pace, naturalness and emotion, with a short note. Then listen to all clips from one line back to back before moving to the next line, which makes differences obvious that a single clip hides.
LINES = [
"Order by October 31 and pay 29.99 dollars. Call 555 0142 for help.",
"Meet the Aurora travel mug, now in three sizes.",
"It keeps drinks hot for twelve hours, fits every cup holder, and survives a drop from the counter.",
"I can't believe it actually survived.",
"Gracias por su compra.",
]
RATES = {"MAI-Voice-2.1-Flash": 15.0, "MAI-Voice-2.1": 22.0, "Sume TTS Router": 47.5}
chars = sum(len(x) for x in LINES)
print("characters:", chars)
for name, per_million in RATES.items():
print("%-22s $%.5f for one pass" % (name, chars * per_million / 1_000_000))
Scoring and deciding
Keep the score sheet small: five lines, four criteria, three engines, sixty cells. Fill it in one sitting, in a fixed order, so fatigue hits every engine equally. Write one sentence per engine on what you would change before shipping. If two engines tie on the sheet, break the tie on price per million characters and on whether the route fits the tools you already use.
Keep the clips. When Microsoft ships the next voice version, or when a competitor announces a price cut, you can replay the same five lines and compare against your saved audio instead of relying on memory. A dated folder of clips is the cheapest regression test a voice project can have.
Finally, note the limits in the sheet: Flash's 45 seconds of audio per generation, and Sume's 20,000 character transcript ceiling per request. A script that is too long for one call needs chunking in either case, and that belongs in the decision.
What it costs to test on Sume
If you want a Sonic comparison point, one pass of the five lines is about 250 characters, which is about $0.012 through Sume's TTS Router at $47.50 per 1M characters. The router takes a model id from its catalog, such as sonic-3.6, a transcript and a voice selector, and returns an async job you poll for a finished file (Jobs and results). I did not read a free listening playground for the router in the docs, so plan on paying cents, not nothing.
Whatever you hear, check the license and disclosure rules for the route you used before you publish synthetic voice, especially in ads. Those are separate from sound quality and often decide which engine you are allowed to ship.
Sources
Related posts
More in Models
- Which AI video models accept an input video on Sume?
Seedance, Wan 3.0, H3, H3 Max and Gemini Omni Flash take video references; Recast, Genjutsu and Omni edit need a source video. Kling and Grok take none.
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
- Which AI video models take a reference audio file on Sume?
Seedance 2.x, Wan 3.0, MiniMax H3 and H3 Max accept a reference audio clip on Sume; Gemini Omni Flash, Kling, Grok, Recast and Genjutsu do not.
- Which Sume image models accept quality? Only five do
Only five Sume image catalog rows list a quality field: GPT Image 2, 2.5 and Sunburst, Ideogram V3 and 4.5. The rest return 400 unsupported_parameter.
Written by Sume