MAI-Voice-2.1-Flash: 150 ms end-to-end and a 45-second audio limit

Microsoft's news post gives MAI-Voice-2.1-Flash 150 ms end-to-end latency and 45 s of audio; the Learn page gives no latency number. What each page supports.

5 min readSume
All posts

MAI-Voice-2.1-Flash has one published latency figure: Microsoft's launch post says it generates "45s of audio" with "end-to-end latency, of a mere 150ms", at $15 per 1M characters. The Learn documentation page describes the model as built for low latency and real-time agents but states no millisecond figure, no per-request audio cap and no price (it links to the Azure Speech pricing page). Read on 2026-10-08.

So the number you can quote is the news post's, and it is a vendor claim. Microsoft's post does not say here what the 150 ms includes, so measure first-byte time from your own region.

What each page says

The news post also claims 55% faster model inference and roughly 60% lower cost than comparable models. The Learn page says MAI-Voice-2.1 (the non-Flash model) prioritizes naturalness and expressivity over latency-critical scenarios, which is Microsoft's own guidance to pick Flash for live calls.

MAI-Voice-2.1-Flash claims by Microsoft page (read 2026-10-08)
ClaimNews postLearn page
Price$15 per 1M charactersNot stated (links to pricing page)
End-to-end latency150 msNot stated
Audio per request45 sNot stated
Languages23 languages, 26 locales23 languages
Preview or SLANot statedPublic preview, no SLA
RegionsNot stated14 regions

What 45 seconds means for scripts

A 45-second cap bounds each synthesis call to a few hundred words. At a typical 150 words per minute (an assumption, not a Microsoft figure), 45 seconds is about 112 words or 650 to 700 characters. Longer scripts have to be split by sentence and the audio joined, which is easy for chat turns and awkward for an audiobook; Microsoft points long-form work at MAI-Voice-2.1 instead.

Where Sume differs

Sume's TTS 1.0 is not a streaming route. The API reference describes it as an async job with polling or a webhook, and says mode: sync is a bounded wait of up to 30 seconds on the HTTP request, not a latency promise. A job can be up to 20,000 characters and 1,200 seconds of audio, so the shape is closer to "render this narration" than "answer this caller".

That makes Sume the wrong tool for a phone agent's first word, and the right one for pre-rendered prompts, ad reads and voice-overs. The latency-by-job-type post sorts which jobs need which.

How to measure it yourself

Time three things separately: the delay from request to first audio byte, the delay to the last byte, and the audio length. A model that reports 150 ms to start but streams slower than real time will still stall a call. Run at least 50 requests at the hours your callers actually use, and report the median and the slowest 5 percent, not the best run.

For Sume's async route the equivalent measurement is the time from submit to a completed job with a downloadable file; the benchmark post shows a Python loop for it. The two numbers answer different questions, so do not put them in one column of a vendor comparison.

Sources

Related posts

More in Models

All Models posts

Written by Sume