Can I use MAI-Voice-2.1 audio commercially? What Microsoft says
Microsoft's MAI voices page says it holds full licensing rights for commercial use, the service is in public preview with no SLA, and cloning is gated.

The Microsoft Learn MAI voices page, dated 2026-10-01 and read on 2026-10-04, says the voices are available to third-party developers and that Microsoft holds full licensing rights for commercial use. That sentence reads as Microsoft's licensing position on the voices; it does not state who owns a given output file, so confirm that point with Microsoft before shipping a paid product.
The same page marks the service as public preview, with no SLA and not recommended for production.
What the page says, line by line
| Topic | What the page says |
|---|---|
| Availability | Available for third-party developers |
| Commercial use | Microsoft holds full licensing rights for commercial use |
| Status | Public preview, no SLA, not recommended for production |
| Calling it | Azure Speech SDK or REST, with SSML and voice names such as en-US-Harper:MAI-Voice-2.1 |
| Cloning | Instant cloning from a 5-60 s clip, gated through the Limited Access Review |
| Pricing | The page points to the Azure Speech pricing page, which I did not read |
What the page does not settle
Nothing on the page tells you who owns a generated clip, whether a customer can resell it as a standalone asset, or how the preview terms change at general availability. Those are contract questions, not documentation ones.
The launch coverage on Unite.ai lists MAI-Voice-2.1 at $22 per 1M characters and the Flash model at $15 per 1M characters. Treat these as reported numbers until you check Azure's own page.
If the answer needs to be firm today
Sume's TTS endpoint is documented in the API reference. It takes a plain transcript and a voice, returns audio, and is priced at $0.0475 per 1,000 characters. Its catalog lists Sonic models, not MAI voices, so it is an alternative, not a MAI wrapper.
- Use MAI voices in preview for prototypes where an SLA is not needed.
- Ask Microsoft in writing about output ownership for paid products.
- Keep consented voices only; cloning is gated for a reason.
- For a production voiceover today, a generally available endpoint avoids the preview caveat.
Pre-launch review list
- Which MAI voice name and region you call.
- Whether you rely on cloning, which needs approval.
- Your own contract wording on output rights.
- A fallback voice if the preview changes.
Sources
Related posts
More in Models
- MAI-Voice-2.1 emotion control vs Sume's emotion field
Microsoft lists emotion control on MAI-Voice-2.1. Sume's TTS takes a free-text emotion string, speed and volume. What each gives you, and how to test it.
- MAI-Voice-2.1 English voices with call-centre or audiobook styles
customer_call_center and audiobook styles sit on en-US Grant, Harper, Sage and en-GB Emily, Harry. Iris, Jasper, Dhruv and Priya are neutral only.
- Which Korean voices does MAI-Voice-2.1 have, and with what styles?
MAI-Voice-2.1 lists four ko-KR voices: Grant, Harper (role styles), Haena (14 emotion styles) and Junho (12). Pair a voice with the right caption style.
- MiniMax H3 Max: the prompt-adherence variant on Sume
fal describes MiniMax H3 Max as tuned for prompt adherence. On Sume, minimax-h3-max runs 480p to 1080p for 5 to 15 s with frames and references.
Written by Sume