MAI-Voice-2.1 neutral-only voices: Hungarian, Romanian, en-IN
Several MAI-Voice-2.1 voices list only a neutral style: five Hungarian, four Romanian and both en-IN voices. Check the style list before you cast an ad.

In Microsoft's MAI-Voice-2.1 list, five of the six Hungarian voices, four of the six Romanian voices and both en-IN voices list only a neutral style, so an upbeat or urgent ad read cannot be styled with mstts:express-as. The catch is not obvious until you open each voice's style row.
Neutral-only voices we found
From the Learn page, read by our count:
| Locale | Neutral-only | Has other styles |
|---|---|---|
| hu-HU | Bence, Grant, Levente, Lilla, Reka | Harper (role styles) |
| ro-RO | Andrei, Elena, Ioana, Radu | Grant (agent, call center, educational, narrator); Harper (audiobook) |
| en-IN | Dhruv, Priya | None |
| pt-PT | Grant, Harper | Rui (emotion styles) |
What to do when the style is not offered
Choose the voice that lists the style you need, or accept a flat read and add energy through script wording and pacing. Do not assume the SSML will error: test a short line first and listen.
Microsoft's own doc sample uses style happiness, while en-US-Harper's list shows happy and joyful, so copy style names from the list, not the sample.
On Sume
Sume TTS has no per-voice style list. The delivery hint is a free-text emotion string up to 64 characters in generation_config (API reference). It is a hint, not a guarantee, so audition it the same way.
Sources
Related posts
More in Models
- MAI-Voice-2.1 has three style lists: emotion, role and expressive
MAI-Voice-2.1 styles come in three vocabularies: 19 emotions, 6 roles and 11 expressive tags. Which voices use the expressive set, and how Sume differs.
- MAI-Voice-2.1 whispering and shouting: the 19 voices that list both
Only 19 MAI-Voice-2.1 voices list whispering and shouting by our count. English-UK and Korean voices do not. Sume has no whisper switch.
- MAI-Voice 50.3% of 4,000 listeners: how to quote the Turing claim
Microsoft said 50.3% of 4,000 listeners rated MAI-Voice as equally or more human-like than human recordings. What it covers, and a cheap test.
- Three characters, three dances: Omni IMAGE_REF and VIDEO_REF tokens
Google's Omni 1.1 demo swaps three dancers for a dog, an octopus and a bear. Here is the same request on Sume, with the 0-based reference tokens in order.
Written by Sume