MAI-Voice-2.1 or Flash for narration? Microsoft's own guidance
Microsoft lists MAI-Voice-2.1 for long-form narration and Flash for agents. The prices, the latency claims and which one to test first for voiceover.

For narration, test MAI-Voice-2.1 first, because Microsoft's docs list it for expressive long-form content, audiobooks, podcasts and voice overs, and describe Flash as the model for real-time agents and call-center flows. Flash costs less, at $15 per million characters against $22.
What Microsoft says about each
From the Microsoft Learn page and the announcement, read on 2026-10-07:
| Point | MAI-Voice-2.1 | MAI-Voice-2.1-Flash |
|---|---|---|
| Best for | Long-form content, audiobooks, podcasts, voice overs | Real-time agents, IVR, call centers |
| Languages | 23 | 23 |
| Price | $22 per 1M characters | $15 per 1M characters |
| Latency claim | Prioritizes expression over latency | 45 seconds of audio with 150 ms end-to-end latency |
| Cloning | Gated instant cloning | Gated instant cloning |
Why Flash is still worth one test
For short ad reads of a few sentences, the expressiveness gap may be small and the savings real. Generate the same three lines on both and listen blind. If listeners cannot tell, take the cheaper model. If the longer read shows drift or flat endings on Flash, move to MAI-Voice-2.1.
Mind the preview label
The docs call the feature a public preview with no SLA. For a client deliverable, keep the approved files and do not rely on regenerating them later.
What Sume adds
Sume's TTS 1.0 charges $0.0475 per 1,000 characters, and an estimate assumes at most 20,000 characters per job. If you want the voice from a Microsoft model, you generate it in Azure and import the file; if you want the audio and the render in one chain, use the Sume job. Check GET /v1/catalog for the live rate.
Sources
Related posts
More in Models
- MAI-Voice-2.1 styles per voice: 21 list one, 18 list 19
Of 97 MAI-Voice-2.1 voices, 21 list one style and 18 list 19. The distribution, and how it shapes casting before a Sume or MAI voiceover.
- MAI-Voice-2.1 has three style lists: emotion, role and expressive
MAI-Voice-2.1 styles come in three vocabularies: 19 emotions, 6 roles and 11 expressive tags. Which voices use the expressive set, and how Sume differs.
- MAI-Voice-2.1 whispering and shouting: the 19 voices that list both
Only 19 MAI-Voice-2.1 voices list whispering and shouting by our count. English-UK and Korean voices do not. Sume has no whisper switch.
- MAI-Voice 50.3% of 4,000 listeners: how to quote the Turing claim
Microsoft said 50.3% of 4,000 listeners rated MAI-Voice as equally or more human-like than human recordings. What it covers, and a cheap test.
Written by Sume