MAI-Voice-2.1 voice cloning is gated; how Sume TTS picks a voice
MAI-Voice-2.1 clones a voice from a 5 to 60 second clip but needs gated access. Sume TTS selects voices by avatar or voice id and has no reference-audio field.

MAI-Voice-2.1 can match a consented voice from a short reference clip, but the feature is gated behind a Microsoft access review. Sume TTS works differently: you choose a voice through an avatar (avatar_id or avatar_handle) or a voice.id, and the TTS request has no reference-audio field. If you need to bring a voice from a recording, Sume's TTS endpoint will not do it; its documented route is selecting an existing voice.
What Microsoft documents
Microsoft Learn says MAI-Voice-2.1 and MAI-Voice-2.1-Flash support instant voice cloning, gated, from a clip of 5 to 60 seconds without extra training. The page says the models have guardrails so that only authorized, consented voices are used, and that you must apply through the Azure Custom Neural Voice and Custom Avatar Limited Access Review, then upload an audio consent and prompt to create a personal voice. Both models are in public preview with no service-level agreement. MAI-Voice-2.1 is $22 per 1M characters and Flash is $15 per 1M characters (Microsoft AI).
What Sume documents
The Sume OpenAPI reference for TTS 1.0 says to provide a transcript plus a voice selector. The selector is the top-level avatar_id or avatar_handle, or voice.id. A voice.id must be a Sume TTS voice UUID or a Voices library id (voi_ plus 32 hex), and other shapes are rejected with 400 invalid_voice_id before a job is queued. An avatar works as a selector exactly when its voice status is ready. The request schema lists no field for a reference clip.
| Question | MAI-Voice-2.1 (Microsoft) | Sume TTS 1.0 |
|---|---|---|
| Clone from a clip | Yes, gated, 5-60 s clip | No reference-audio field in the request |
| How a voice is chosen | Voice name such as en-US-Harper:MAI-Voice-2.1 in SSML | avatar_id, avatar_handle or voice.id |
| Price | $22 per 1M characters | $0.0475 per 1,000 characters |
| Status | Public preview, no SLA | Listed in the public catalog |
Choosing
If your product is a branded voice built from a consented recording, MAI's gated route is the documented path on this comparison, and you carry the consent paperwork. If your product is video and you want one stable voice per character, Sume's avatar route gives you a handle to reuse across scripts, and the same voice can be used for TTS and for talking clips. Creating an avatar is $0.95 per avatar in the catalog. Whatever you choose, store the consent and the voice id with the project, and test a sample before a large batch.
Cost of the same read
For a 1,000 character script, MAI-Voice-2.1 at $22 per 1M characters is $0.022, MAI-Voice-2.1-Flash at $15 per 1M is $0.015, and Sume TTS at $0.0475 per 1,000 is $0.0475. Sume's rate is roughly twice MAI-Voice-2.1's, and the difference is the price of a job inside one API with video, captions and a render step. If you already run Azure and only need speech, the Microsoft rate is lower.
Sources
Related posts
More in Comparisons
- MAI-Voice-2.1 emotion control vs Sume TTS emotion, speed and volume
MAI-Voice-2.1 lists emotion control. Sume TTS takes generation_config: emotion (1-64 chars), speed 0.6-1.5, volume 0.5-2. A three-take test costs 3 cents.
- MAI-Voice-2.1 zero-shot voice prompting vs a Sume voice id or avatar
MAI-Voice-2.1 lists zero-shot voice prompting. Sume TTS takes a voice id or an avatar's voice, not a sample clip. What each request can and cannot express.
- MiniMax H3 768p vs Gemini Omni Flash 720p: an 8-second clip
On Sume, an 8-second clip is $0.60 on minimax-h3 at 768p and $1.00 on gemini-omni-flash-1.1 at 720p, with $0.30 at Omni's 360p. Price and constraints.
- MiniMax H3 Max vs Wan 3.0 at 1080p: 5, 10 and 15 seconds on Sume
On Sume, MiniMax H3 Max at 1080p costs $0.20 a second and Wan 3.0 costs $0.25. For 5, 10 and 15 seconds: $1.00/$2.00/$3.00 vs $1.25/$2.50/$3.75.
Written by Sume