MAI-Voice-2.1-Flash: 150ms for 45 seconds of audio, for batch TTS

Microsoft says MAI-Voice-2.1-Flash makes 45s of audio at 150ms end-to-end latency, at $15 per 1M characters. What that does and does not tell a batch TTS user.

4 min readSume
All posts

For batch text to speech, the headline number matters less than it looks. Microsoft's Oct 1, 2026 post says MAI-Voice-2.1-Flash produces 45 seconds of audio at 150ms end-to-end latency, with 55% faster inference and roughly 60% lower cost, at $15 per million characters. The page, as I read it, does not say whether 150ms is time to the first audio or time to the finished 45 seconds, so I would not treat it as a throughput figure. For a batch job the figures that matter are cost per character, per-request limits and queue behaviour, and the post gives only the first.

What is stated

Microsoft AI news post dated Oct 1, 2026, read 2026-10-05; all figures are vendor claims
ItemMAI-Voice-2.1MAI-Voice-2.1-Flash
Price$22 per 1M characters$15 per 1M characters
Latency claimNot stated for the standard model in the text I read45s of audio at 150ms end-to-end
Speed and cost claim-55% faster inference, about 60% cheaper
Languages23 languages, 26 localesNot itemised separately in the text I read
Where to get itOpenRouter, Microsoft Foundry, MAI Playground, Vercel (LiveKit coming soon)Same

Batch cost, worked out

Using the listed prices: a 20,000-character script is 0.02 x $15 = $0.30 on Flash and 0.02 x $22 = $0.44 on MAI-Voice-2.1. A 1,000-character line is $0.015 on Flash. The post's 60% cheaper claim does not match the $15 vs $22 gap (32%), so the baseline for that claim is not these two prices.

For comparison, Sume's TTS Router bills Cartesia Sonic at list x1.25, which is $0.0475 per 1K characters, so the same 20,000-character job is $0.95 before the 5.5% platform fee. Sume does not ship MAI-Voice, and the Sume router is a job API with no streaming in v1, which suits produced audio.

Questions to ask before you switch

  • Is 150ms measured to first byte or to the last of the 45 seconds? That changes whether it is a streaming figure or a batch one.
  • What is the per-request character limit? The text I read does not state one.
  • Which of the 23 languages and 26 locales does your catalog need, and has a native listener checked them?
  • Which route are you using (OpenRouter, Foundry, Vercel)? Each may add its own billing and rate limits.

Sources

Related posts

More in Models

All Models posts

Written by Sume