Where to try MAI-Voice-2.1 before paying: a 10-minute listening test

Microsoft lists the MAI Playground, Copilot Audio Expressions and Foundry for MAI-Voice-2.1. Run a ten-minute listening test with a fixed script and cost it.

5 min readSume
All posts

The fastest free-to-low-cost way to hear MAI-Voice-2.1 is Microsoft's own surfaces. The model page says it is available in the MAI Playground, in Copilot Audio Expressions and in Microsoft Foundry (Azure Speech), and the launch post adds Vercel and OpenRouter (read 2026-10-03). Use the same fixed test script on every engine you consider, because a model that sounds great reading its own demo text can stumble on your product name and your prices.

This post is the script and the scoring, not a review. Listen yourself, and keep the sheet.

Access routes

The model page lists about 550 ms model inference for MAI-Voice-2.1 and about 45 ms for Flash. For a listening test you do not need latency; you need a script that exposes the weak spots.

From Microsoft AI's MAI-Voice-2.1 page and launch post, read 2026-10-03. Prices are Microsoft's list prices per 1M characters.
RouteBest forNotes
MAI PlaygroundQuick listeningNo integration needed
Copilot Audio ExpressionsSeeing the emotion controlsNamed on the model page
Microsoft FoundryProduction use inside Azure$22 per 1M (2.1), $15 per 1M (Flash)
Vercel, OpenRouterCalling it from an existing gatewayCheck the gateway's own rate

A fixed ten-minute test script

Use five lines, each written to break something. A line with numbers and units: a price, a phone number, a date. A line with your brand and product names. A long sentence with two clauses and a list. A short, emotional line the voice should act. And a line in a second language if you need one. Record the same five lines on each engine in the same order.

Score each clip 1 to 5 on pronunciation, pace, naturalness and emotion, with a short note. Then listen to all clips from one line back to back before moving to the next line, which makes differences obvious that a single clip hides.

LINES = [
    "Order by October 31 and pay 29.99 dollars. Call 555 0142 for help.",
    "Meet the Aurora travel mug, now in three sizes.",
    "It keeps drinks hot for twelve hours, fits every cup holder, and survives a drop from the counter.",
    "I can't believe it actually survived.",
    "Gracias por su compra.",
]
RATES = {"MAI-Voice-2.1-Flash": 15.0, "MAI-Voice-2.1": 22.0, "Sume TTS Router": 47.5}

chars = sum(len(x) for x in LINES)
print("characters:", chars)
for name, per_million in RATES.items():
    print("%-22s $%.5f for one pass" % (name, chars * per_million / 1_000_000))

Scoring and deciding

Keep the score sheet small: five lines, four criteria, three engines, sixty cells. Fill it in one sitting, in a fixed order, so fatigue hits every engine equally. Write one sentence per engine on what you would change before shipping. If two engines tie on the sheet, break the tie on price per million characters and on whether the route fits the tools you already use.

Keep the clips. When Microsoft ships the next voice version, or when a competitor announces a price cut, you can replay the same five lines and compare against your saved audio instead of relying on memory. A dated folder of clips is the cheapest regression test a voice project can have.

Finally, note the limits in the sheet: Flash's 45 seconds of audio per generation, and Sume's 20,000 character transcript ceiling per request. A script that is too long for one call needs chunking in either case, and that belongs in the decision.

What it costs to test on Sume

If you want a Sonic comparison point, one pass of the five lines is about 250 characters, which is about $0.012 through Sume's TTS Router at $47.50 per 1M characters. The router takes a model id from its catalog, such as sonic-3.6, a transcript and a voice selector, and returns an async job you poll for a finished file (Jobs and results). I did not read a free listening playground for the router in the docs, so plan on paying cents, not nothing.

Whatever you hear, check the license and disclosure rules for the route you used before you publish synthetic voice, especially in ads. Those are separate from sound quality and often decide which engine you are allowed to ship.

Sources

Related posts

More in Models

All Models posts

Written by Sume