MAI-Voice-2.1 voice ID format and which voices have styles
A MAI-Voice voice is named like en-US-Harper:MAI-Voice-2.1-Flash. Style lists differ by voice: some have 20 styles, some only neutral.

To pick a MAI-Voice-2.1 voice, write the full voice ID plus a model suffix: en-US-Harper:MAI-Voice-2.1 or en-US-Harper:MAI-Voice-2.1-Flash. The Microsoft Learn page says the same suffix applies to any supported prebuilt voice, and that every managed voice works with both models.
The part that trips up first builds is styles. The supported style list is per voice, not per model, and it varies a lot.
The ID and the style tag
A voice ID has three parts: locale, a first name, then the colon-separated model. Styles are set with the SSML mstts:express-as element and its style attribute. Microsoft's example uses style="happiness" with en-US-Harper, whose list on the page does not contain that exact word: it lists happy and joyful. Check each voice's list rather than copying the sample blindly.
| Voice ID | Gender | Styles listed |
|---|---|---|
| en-US-Harper | Female | 20, including agent, audiobook, narrator, whispering |
| en-US-Olivia | Female | 19, including excited, sad, shouting, whispering |
| en-US-Grant | Male | 6: agent, audiobook, customer_call_center, educational, narrator, neutral |
| en-US-Sage | Male | 6: same six as Grant |
| en-US-Iris | Female | 1: neutral |
Two families of style names
The page shows two vocabularies. Some voices list role styles such as agent, audiobook, customer_call_center, educational and narrator. Others list emotion styles such as angry, excited, sad and whispering. A few Spanish, Dutch, Russian, Thai and Turkish voices list a third set, for example adventurous, caringempathy and friendlycheerful.
The consequence: you cannot write one SSML template that sets one style across every voice. A template for a call-center agent should use a role style, and a template for a dramatic read should use an emotion style, each on a voice that lists it.
A safe way to build the template
Keep a small table in your own code that maps voice ID to the styles you have confirmed, and refuse a style that is not in it. That turns a silent fallback into a visible error in testing.
- Store the voice ID with the suffix as one string, so Flash and the full model are separate entries.
- Store the confirmed style list next to it and the date you read it.
- Reject unknown styles before you call the API.
- Re-read the page on each release, since Microsoft says it adds locales and managed voices over time.
The same question on Sume
A Sume TTS request takes a plain transcript and has no SSML field. A voice is chosen by an avatar reference, or by a voice id that is either a UUID or a Voices-library id starting voi_; any other shape is rejected with a 400 before a job is queued and before credits are reserved, per the Sume OpenAPI spec. Delivery controls live in generation_config: a volume from 0.5 to 2, a speed from 0.6 to 1.5, and a free-text emotion guide of up to 64 characters. There is no per-voice style table to check, which is simpler, but also less explicit about what a voice will do.
Sources
Related posts
More in Models
- MAI-Voice instant cloning: gated access and a 5-60 second clip
MAI-Voice cloning needs approval through Microsoft's Limited Access Review and a 5-60 second consented clip. What that means for your launch plan.
- Make AI Video Follow Your Audio: Which Models Take Audio References
Seedance, Wan 3.0 and MiniMax H3 accept an audio reference on Sume; Omni does not. Request shape, Wan limits and what an audio reference is not.
- Mercury Voice is enterprise-only: pricing and what to ask sales
Mercury Voice is GA for enterprise customers only. List $0.40/$1.50 per million tokens, launch price $0.20/$0.75. Questions to ask before you commit.
- Mercury Voice p95 750 ms: one turn in twenty is slower
Inception reports a 320 ms median and 750 ms p95 for Mercury Voice. At p95, about one turn in 20 is slower. What that means over a 10-turn call.
Written by Sume