MAI-Voice-2.1 voice cloning is gated; how Sume TTS picks a voice

MAI-Voice-2.1 clones a voice from a 5 to 60 second clip but needs gated access. Sume TTS selects voices by avatar or voice id and has no reference-audio field.

5 min readSume
All posts

MAI-Voice-2.1 can match a consented voice from a short reference clip, but the feature is gated behind a Microsoft access review. Sume TTS works differently: you choose a voice through an avatar (avatar_id or avatar_handle) or a voice.id, and the TTS request has no reference-audio field. If you need to bring a voice from a recording, Sume's TTS endpoint will not do it; its documented route is selecting an existing voice.

What Microsoft documents

Microsoft Learn says MAI-Voice-2.1 and MAI-Voice-2.1-Flash support instant voice cloning, gated, from a clip of 5 to 60 seconds without extra training. The page says the models have guardrails so that only authorized, consented voices are used, and that you must apply through the Azure Custom Neural Voice and Custom Avatar Limited Access Review, then upload an audio consent and prompt to create a personal voice. Both models are in public preview with no service-level agreement. MAI-Voice-2.1 is $22 per 1M characters and Flash is $15 per 1M characters (Microsoft AI).

What Sume documents

The Sume OpenAPI reference for TTS 1.0 says to provide a transcript plus a voice selector. The selector is the top-level avatar_id or avatar_handle, or voice.id. A voice.id must be a Sume TTS voice UUID or a Voices library id (voi_ plus 32 hex), and other shapes are rejected with 400 invalid_voice_id before a job is queued. An avatar works as a selector exactly when its voice status is ready. The request schema lists no field for a reference clip.

Voice selection and cloning, from Microsoft Learn and Microsoft AI (read 2026-10-09) and the Sume OpenAPI and catalog (read 2026-10-09).
QuestionMAI-Voice-2.1 (Microsoft)Sume TTS 1.0
Clone from a clipYes, gated, 5-60 s clipNo reference-audio field in the request
How a voice is chosenVoice name such as en-US-Harper:MAI-Voice-2.1 in SSMLavatar_id, avatar_handle or voice.id
Price$22 per 1M characters$0.0475 per 1,000 characters
StatusPublic preview, no SLAListed in the public catalog

Choosing

If your product is a branded voice built from a consented recording, MAI's gated route is the documented path on this comparison, and you carry the consent paperwork. If your product is video and you want one stable voice per character, Sume's avatar route gives you a handle to reuse across scripts, and the same voice can be used for TTS and for talking clips. Creating an avatar is $0.95 per avatar in the catalog. Whatever you choose, store the consent and the voice id with the project, and test a sample before a large batch.

Cost of the same read

For a 1,000 character script, MAI-Voice-2.1 at $22 per 1M characters is $0.022, MAI-Voice-2.1-Flash at $15 per 1M is $0.015, and Sume TTS at $0.0475 per 1,000 is $0.0475. Sume's rate is roughly twice MAI-Voice-2.1's, and the difference is the price of a job inside one API with video, captions and a render step. If you already run Azure and only need speech, the Microsoft rate is lower.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume