MAI-Voice-2.1 styles per voice: 21 list one, 18 list 19

Of 97 MAI-Voice-2.1 voices, 21 list one style and 18 list 19. The distribution, and how it shapes casting before a Sume or MAI voiceover.

5 min readSume
All posts

The 97 MAI-Voice-2.1 voices do not have comparable style lists. I counted the voice table: 21 voices list exactly one style (neutral), 21 list the six role styles, and 18 list 19. Only one voice, en-US Harper, lists 20 (read 2026-10-07).

The distribution

Counting the style column on Microsoft Learn gives 13 distinct list lengths. The two biggest groups sit at the ends: 21 voices with one style, and 18 with 19. Together with the 21 role voices, three groups hold 60 of 97 voices.

Voices by number of listed styles (read 2026-10-07)
Styles listedVoices
121
24
34
44
56
621
119
122
131
145
171
1918
201

What the groups mean

The 21 single-style voices are neutral-only, so any emotion has to come from the script. The 21 six-style voices have the role set (agent, audiobook, customer_call_center, educational, narrator, neutral), with no emotion styles. The 19-style group has the long emotion list plus whispering, shouting and softvoice. Eleven-style voices (nine of them) have a shorter emotion list.

If you shop by locale, check the group, not the label. Two voices in one locale can differ by 18 styles.

  • Neutral only: 21 voices.
  • Role styles: 21 voices with six items.
  • Expressive: 18 voices with 19 items, plus one with 20.

Using the count

A review that says the model supports emotion is true of fewer than half the voices. Count the expressive voices, those with at least 11 styles: 9 + 2 + 1 + 5 + 1 + 18 + 1 = 37, which is 38 percent of the list.

The cost does not vary by voice. Microsoft lists $22 per million characters for MAI-Voice-2.1 and $15 for Flash, whatever voice is named.

On Sume

Sume has no per-voice style lists. Its TTS call has an emotion string (up to 64 characters), speed and volume that apply to any voice, at $0.0475 per 1,000 characters. If your priority is a consistent control surface across every voice in the library, that is the simpler model. If your priority is a named, labelled style on a specific voice, the Microsoft table is richer.

Try both on a 40-word line and decide by ear, with the same script and no post-processing.

How to use this table

Before casting, filter the voice table by locale, then by the style group you need, and only then by gender and sample. Doing it in that order avoids falling for a sample of a voice that cannot do what the script needs.

Two practical filters follow from the table. If the script needs more than a neutral read, drop the 42 voices that list one or six styles, which leaves 55 candidates. If it also needs a named role such as customer_call_center, the candidate set is the role voices only, and emotion is out of reach.

Keep the count current. The list is for a model in public preview, and a new voice or a revised style list changes the groups above. Re-count whenever Microsoft updates the page, and note the read date in your own sheet.

Sources

Related posts

More in Models

All Models posts

Written by Sume