MAI-Voice-2.1 has three style lists: emotion, role and expressive
MAI-Voice-2.1 styles come in three vocabularies: 19 emotions, 6 roles and 11 expressive tags. Which voices use the expressive set, and how Sume differs.

Microsoft's MAI-Voice-2.1 page uses three separate style vocabularies: a 19-item emotion set, a 6-item role set and an 11-item expressive set, and each voice supports only some of them. The expressive set appears only on eight voices in Spanish (Spain), Dutch, Russian, Thai and Turkish.
The three sets
Names come from the Learn page.
| Set | Count | Examples |
|---|---|---|
| Emotion | 19 | angry, excited, happy, sad, softvoice, whispering, shouting |
| Role | 6 | agent, audiobook, customer_call_center, educational, narrator, neutral |
| Expressive | 11 | adventurous, caringempathy, curious, encouraging, friendlycheerful, nostalgic, serious |
Expressive voices
Our count lists these as expressive-set voices: es-ES-Marta, nl-NL-Sander, ru-RU-Lev, ru-RU-Masha, th-TH-Krit, th-TH-Nattapong, tr-TR-Aydin and tr-TR-Elif. A script written with emotion names will not map onto them, so keep a style sheet per voice.
Why a single style sheet breaks
If your creative brief says excited, that word exists in the emotion and expressive sets, but happy exists only in the emotion set. A brief that names customer_call_center only applies to a few en and ro voices. Write the brief in plain language and translate it per voice.
Sume avoids the problem by using one free-text emotion string (64 characters) instead of a controlled vocabulary, per the API reference. The tradeoff is that nothing validates your wording.
Sources
Related posts
More in Models
- MAI-Voice-2.1 whispering and shouting: the 19 voices that list both
Only 19 MAI-Voice-2.1 voices list whispering and shouting by our count. English-UK and Korean voices do not. Sume has no whisper switch.
- MAI-Voice 50.3% of 4,000 listeners: how to quote the Turing claim
Microsoft said 50.3% of 4,000 listeners rated MAI-Voice as equally or more human-like than human recordings. What it covers, and a cheap test.
- Three characters, three dances: Omni IMAGE_REF and VIDEO_REF tokens
Google's Omni 1.1 demo swaps three dancers for a dog, an octopus and a bear. Here is the same request on Sume, with the 0-based reference tokens in order.
- MiniMax H3 outputs 24 fps and 32 kHz stereo: what Sume's H3 adds
MiniMax's H3 open-source post lists 24 fps video and 32 kHz stereo audio, 4 to 15 s clips and 768p by default. Sume's minimax-h3 takes 5 to 15 s.
Written by Sume