MAI-Voice-2.1-Flash: 150ms for 45 seconds of audio, for batch TTS
Microsoft says MAI-Voice-2.1-Flash makes 45s of audio at 150ms end-to-end latency, at $15 per 1M characters. What that does and does not tell a batch TTS user.

For batch text to speech, the headline number matters less than it looks. Microsoft's Oct 1, 2026 post says MAI-Voice-2.1-Flash produces 45 seconds of audio at 150ms end-to-end latency, with 55% faster inference and roughly 60% lower cost, at $15 per million characters. The page, as I read it, does not say whether 150ms is time to the first audio or time to the finished 45 seconds, so I would not treat it as a throughput figure. For a batch job the figures that matter are cost per character, per-request limits and queue behaviour, and the post gives only the first.
What is stated
| Item | MAI-Voice-2.1 | MAI-Voice-2.1-Flash |
|---|---|---|
| Price | $22 per 1M characters | $15 per 1M characters |
| Latency claim | Not stated for the standard model in the text I read | 45s of audio at 150ms end-to-end |
| Speed and cost claim | - | 55% faster inference, about 60% cheaper |
| Languages | 23 languages, 26 locales | Not itemised separately in the text I read |
| Where to get it | OpenRouter, Microsoft Foundry, MAI Playground, Vercel (LiveKit coming soon) | Same |
Batch cost, worked out
Using the listed prices: a 20,000-character script is 0.02 x $15 = $0.30 on Flash and 0.02 x $22 = $0.44 on MAI-Voice-2.1. A 1,000-character line is $0.015 on Flash. The post's 60% cheaper claim does not match the $15 vs $22 gap (32%), so the baseline for that claim is not these two prices.
For comparison, Sume's TTS Router bills Cartesia Sonic at list x1.25, which is $0.0475 per 1K characters, so the same 20,000-character job is $0.95 before the 5.5% platform fee. Sume does not ship MAI-Voice, and the Sume router is a job API with no streaming in v1, which suits produced audio.
Questions to ask before you switch
- Is 150ms measured to first byte or to the last of the 45 seconds? That changes whether it is a streaming figure or a batch one.
- What is the per-request character limit? The text I read does not state one.
- Which of the 23 languages and 26 locales does your catalog need, and has a native listener checked them?
- Which route are you using (OpenRouter, Foundry, Vercel)? Each may add its own billing and rate limits.
Sources
Related posts
More in Models
- MAI-Voice-2.1-Flash at 45 ms: does a rendered avatar need fast TTS?
Microsoft lists MAI-Voice-2.1-Flash at about 45 ms of inference. A rendered avatar clip does not benefit from it. Where the latency shows up in a Sume job.
- MAI-Voice-2.1 Hindi: six voices listed, and a Hindi request on Sume
Microsoft's MAI-Voice-2.1 page lists six hi-IN voices with different style sets. Compare that with a Hindi request on Sume TTS and what the language field does.
- Mandarin Chinese text to speech API: MAI-Voice-2.1 zh-CN vs Sume zh
MAI-Voice-2.1 lists Chinese (Simplified) as zh-CN; Sume tags zh and bills by character. Audition Mandarin ad copy for 1 cent and know what to check.
- Which Foundry models add C2PA and a watermark: GPT Image, FLUX.2, MAI
Microsoft's Foundry page lists the image, audio and voice models that get C2PA credentials and watermarks. What that does and does not tell you about Sume.
Written by Sume