MAI-Voice-2.1 or Flash for narration? Microsoft's own guidance

Microsoft lists MAI-Voice-2.1 for long-form narration and Flash for agents. The prices, the latency claims and which one to test first for voiceover.

5 min readSume
All posts

For narration, test MAI-Voice-2.1 first, because Microsoft's docs list it for expressive long-form content, audiobooks, podcasts and voice overs, and describe Flash as the model for real-time agents and call-center flows. Flash costs less, at $15 per million characters against $22.

What Microsoft says about each

From the Microsoft Learn page and the announcement, read on 2026-10-07:

MAI-Voice-2.1 and Flash (read 2026-10-07)
PointMAI-Voice-2.1MAI-Voice-2.1-Flash
Best forLong-form content, audiobooks, podcasts, voice oversReal-time agents, IVR, call centers
Languages2323
Price$22 per 1M characters$15 per 1M characters
Latency claimPrioritizes expression over latency45 seconds of audio with 150 ms end-to-end latency
CloningGated instant cloningGated instant cloning

Why Flash is still worth one test

For short ad reads of a few sentences, the expressiveness gap may be small and the savings real. Generate the same three lines on both and listen blind. If listeners cannot tell, take the cheaper model. If the longer read shows drift or flat endings on Flash, move to MAI-Voice-2.1.

Mind the preview label

The docs call the feature a public preview with no SLA. For a client deliverable, keep the approved files and do not rely on regenerating them later.

What Sume adds

Sume's TTS 1.0 charges $0.0475 per 1,000 characters, and an estimate assumes at most 20,000 characters per job. If you want the voice from a Microsoft model, you generate it in Azure and import the file; if you want the audio and the render in one chain, use the Sume job. Check GET /v1/catalog for the live rate.

Sources

Related posts

More in Models

All Models posts

Written by Sume