MAI-Voice-2.1 voice ID format and which voices have styles

A MAI-Voice voice is named like en-US-Harper:MAI-Voice-2.1-Flash. Style lists differ by voice: some have 20 styles, some only neutral.

5 min readSume
All posts

To pick a MAI-Voice-2.1 voice, write the full voice ID plus a model suffix: en-US-Harper:MAI-Voice-2.1 or en-US-Harper:MAI-Voice-2.1-Flash. The Microsoft Learn page says the same suffix applies to any supported prebuilt voice, and that every managed voice works with both models.

The part that trips up first builds is styles. The supported style list is per voice, not per model, and it varies a lot.

The ID and the style tag

A voice ID has three parts: locale, a first name, then the colon-separated model. Styles are set with the SSML mstts:express-as element and its style attribute. Microsoft's example uses style="happiness" with en-US-Harper, whose list on the page does not contain that exact word: it lists happy and joyful. Check each voice's list rather than copying the sample blindly.

Style coverage of five en-US voices on Microsoft Learn (read 2026-10-03)
Voice IDGenderStyles listed
en-US-HarperFemale20, including agent, audiobook, narrator, whispering
en-US-OliviaFemale19, including excited, sad, shouting, whispering
en-US-GrantMale6: agent, audiobook, customer_call_center, educational, narrator, neutral
en-US-SageMale6: same six as Grant
en-US-IrisFemale1: neutral

Two families of style names

The page shows two vocabularies. Some voices list role styles such as agent, audiobook, customer_call_center, educational and narrator. Others list emotion styles such as angry, excited, sad and whispering. A few Spanish, Dutch, Russian, Thai and Turkish voices list a third set, for example adventurous, caringempathy and friendlycheerful.

The consequence: you cannot write one SSML template that sets one style across every voice. A template for a call-center agent should use a role style, and a template for a dramatic read should use an emotion style, each on a voice that lists it.

A safe way to build the template

Keep a small table in your own code that maps voice ID to the styles you have confirmed, and refuse a style that is not in it. That turns a silent fallback into a visible error in testing.

  • Store the voice ID with the suffix as one string, so Flash and the full model are separate entries.
  • Store the confirmed style list next to it and the date you read it.
  • Reject unknown styles before you call the API.
  • Re-read the page on each release, since Microsoft says it adds locales and managed voices over time.

The same question on Sume

A Sume TTS request takes a plain transcript and has no SSML field. A voice is chosen by an avatar reference, or by a voice id that is either a UUID or a Voices-library id starting voi_; any other shape is rejected with a 400 before a job is queued and before credits are reserved, per the Sume OpenAPI spec. Delivery controls live in generation_config: a volume from 0.5 to 2, a speed from 0.6 to 1.5, and a free-text emotion guide of up to 64 characters. There is no per-voice style table to check, which is simpler, but also less explicit about what a voice will do.

Sources

Related posts

More in Models

All Models posts

Written by Sume