MAI-Voice-2.1 English voices with call-centre or audiobook styles
customer_call_center and audiobook styles sit on en-US Grant, Harper, Sage and en-GB Emily, Harry. Iris, Jasper, Dhruv and Priya are neutral only.

Per the Microsoft Learn MAI voices page read on 2026-10-04, the customer_call_center and audiobook styles are on en-US Grant, Harper and Sage and on en-GB Emily and Harry. Iris, Jasper, Dhruv and Priya offer neutral only, and Ethan, Olivia and Isla list emotion styles instead of roles.
The English list
| Voice | Locale | Style type |
|---|---|---|
| Ethan | en-US | 19 emotion styles |
| Olivia | en-US | 19 emotion styles |
| Grant | en-US | 6 role styles (agent, audiobook, customer_call_center, educational, narrator, neutral) |
| Sage | en-US | Same 6 role styles |
| Harper | en-US | 20 styles mixing roles and emotions |
| Iris | en-US | Neutral only |
| Jasper | en-US | Neutral only |
| Emily | en-GB | Role styles and some emotions |
| Harry | en-GB | Role styles and some emotions |
| Isla | en-AU | 19 emotion styles |
| Dhruv | en-IN | Neutral only |
| Priya | en-IN | Neutral only |
What this means for an IVR or a book
For an IVR line pick a voice that lists customer_call_center. For long-form narration, pick one that lists audiobook or narrator. Remember the out-of-scope note on the page: the long-form model prioritises naturalness over latency, so it is not the choice for live conversation.
The voice name is a full id, for example en-US-Harper:MAI-Voice-2.1, and styles are applied through SSML on the Azure Speech SDK or REST.
A Sume alternative for scripted lines
A voicemail or IVR prompt is a short scripted line, so character count sets the cost. At Sume's $0.0475 per 1,000 characters, a 400-character prompt is $0.019. The Sume endpoint has no SSML or style field; you give a plain transcript and a voice.
- 400 characters x $0.0475 / 1,000 = $0.019.
- 20 prompts of that size = $0.38.
- Re-record a single line without paying for the rest.
Caveat
MAI voices are in public preview with no SLA, per the same page, so keep a fallback for anything customer-facing.
Sources
Related posts
More in Models
- Which Korean voices does MAI-Voice-2.1 have, and with what styles?
MAI-Voice-2.1 lists four ko-KR voices: Grant, Harper (role styles), Haena (14 emotion styles) and Junho (12). Pair a voice with the right caption style.
- MiniMax H3 Max: the prompt-adherence variant on Sume
fal describes MiniMax H3 Max as tuned for prompt adherence. On Sume, minimax-h3-max runs 480p to 1080p for 5 to 15 s with frames and references.
- MiniMax H3 limits: 9 images, 3 videos, 3 audio, file caps
MiniMax's H3 guide caps prompts at 7,000 characters and references at 9 images, 3 videos and 3 audio files. Cheat sheet with the Sume limits beside it.
- MiniMax H3 references on Sume: 9 images, 3 videos, 3 audio, 12 total
minimax-h3 and minimax-h3-max accept 9 images, 3 videos and 3 audio files, 12 in total, and audio cannot be the only reference. Duration rules and errors.
Written by Sume