MAI-Voice instant cloning: gated access and a 5-60 second clip
MAI-Voice cloning needs approval through Microsoft's Limited Access Review and a 5-60 second consented clip. What that means for your launch plan.

Can you clone a voice with MAI-Voice-2.1 today? Only after approval: Microsoft's Learn page says instant voice cloning is gated, needs an application through the Custom Neural Voice and Custom Avatar Limited Access Review, and works from a short clip, recommended at 5 to 60 seconds, with no training.
The page adds that only authorized, licensed voices can be synthesized in production and that no unlicensed voice cloning is possible. So the gate is part of the product, not a formality you can plan around.
The four steps Microsoft lists
The Learn page gives a fixed sequence for custom voices. Each step is a dependency in your schedule.
- Apply through the Limited Access Review, which is a separate approval from getting a Speech resource.
- After approval, use the personal voice APIs in the Speech SDK custom-voice samples.
- Upload an audio consent statement and the prompt audio to create a personal voice.
- Synthesize with a speaker profile ID inside an
mstts:ttsembeddingelement, using the MAI-Voice model name.
How it compares with two other routes
Cloning appears across the voice tools in this week's launches, but the access terms differ. The Pocket TTS repository accepts plain WAV files as the voice input and carries an MIT license, so there is no approval step in the software, and the README prohibits cloning without explicit and lawful consent, which leaves checking consent and rights to your own process. Sume's route uses workspace avatars or the Voices library: you pass an avatar reference or a voice id, and Sume resolves the voice at submit time.
| Tool | How a custom voice enters | Gate |
|---|---|---|
| MAI-Voice-2.1 and Flash | Reference clip, recommended 5-60 seconds, speaker profile ID | Limited Access Review, consent upload |
| Kyutai Pocket TTS | Plain WAV file as the audio prompt | No technical gate; MIT license, and the README prohibits cloning without explicit and lawful consent |
| Sume TTS 1.0 | avatar_id, avatar_handle, or voice.id (UUID or voi_ id) | Workspace access; the voice must belong to your workspace avatar or library |
What to keep on file whichever route you take
A gate from the vendor does not replace your own record of who agreed to what. Keep the speaker's written consent, the scope of use, the date, and the clip you used. The consent-record checklist lists the fields.
Plan for the review time as a dependency. If a campaign voice must be cloned by a date, apply before you write the script, because the approval, not the audio generation, is the slow step.
If you only need a consistent branded narrator and not the exact voice of a person, a prebuilt voice avoids the gate entirely. MAI-Voice-2.1 ships licensed curated voices across its supported languages for that purpose.
Finally, check the language you need. The cloned voice carries across the supported languages according to Microsoft, but your approval, consent statement and licence terms should name every language and every channel where the cloned voice will be heard.
Sources
Related posts
More in Models
- Make AI Video Follow Your Audio: Which Models Take Audio References
Seedance, Wan 3.0 and MiniMax H3 accept an audio reference on Sume; Omni does not. Request shape, Wan limits and what an audio reference is not.
- Mercury Voice is enterprise-only: pricing and what to ask sales
Mercury Voice is GA for enterprise customers only. List $0.40/$1.50 per million tokens, launch price $0.20/$0.75. Questions to ask before you commit.
- Mercury Voice p95 750 ms: one turn in twenty is slower
Inception reports a 320 ms median and 750 ms p95 for Mercury Voice. At p95, about one turn in 20 is slower. What that means over a 10-turn call.
- MiniMax H3 2K and 4K upscale on Sume: H3 accepts them, H3 Max does not
minimax-h3 will price a 2K or 4K request even though its resolution list shows only 480p and 768p. minimax-h3-max rejects both. Costs for 5 to 15 seconds.
Written by Sume