MAI custom voice SSML (ttsembedding) vs Sume voice ids: what is gated

Microsoft's MAI-Voice-2.1 custom voice sits behind Limited Access Review and uses a speakerProfileId in SSML. Sume picks voices by id, with no clone widget.

4 min readSume
All posts

To use a custom MAI-Voice-2.1 voice you need Microsoft's Limited Access Review approval, then reference a speaker profile in SSML with a ttsembedding element. Sume TTS selects voices by id and its MCP tool says there is no voice-cloning or voice-upload widget.

How Microsoft's flow works

Per the Learn page: a reference clip of 5 to 60 seconds, consent audio and a prompt are uploaded, and the SSML names the base model and the profile (<voice name='MAI-Voice-2.1'> wrapping <mstts:ttsembedding speakerProfileId='...'>). Access is gated, so plan lead time.

Custom voice path, read 2026-10-07
StepMAI-Voice-2.1Sume TTS 1.0
AccessLimited Access ReviewNone for voice selection
Voice referencespeakerProfileId in SSMLvoice.id (UUID or voi_ + 32 hex)
Clone from a clipYes, 5 to 60 s referenceNo widget in the TTS tool
MarkupSSMLNo SSML

What Sume does instead

Sume validates voice.id (a bad value returns 400 invalid_voice_id) or takes an avatar voice marked ready. Voices are created through the voices tooling rather than by uploading a sample inside the TTS call (API reference).

If your brief requires the voice of a specific real person, Sume's TTS endpoint is not the path, and we are not claiming it is. Get consent and use a vendor flow built for it.

Pricing context

Microsoft lists $22 per million characters for standard and $15 for Flash; whether custom voices use the same rate is not stated on the pages we read.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume