MAI-Voice-2.1-Flash: 150 ms end-to-end and a 45-second audio limit
Microsoft's news post gives MAI-Voice-2.1-Flash 150 ms end-to-end latency and 45 s of audio; the Learn page gives no latency number. What each page supports.

MAI-Voice-2.1-Flash has one published latency figure: Microsoft's launch post says it generates "45s of audio" with "end-to-end latency, of a mere 150ms", at $15 per 1M characters. The Learn documentation page describes the model as built for low latency and real-time agents but states no millisecond figure, no per-request audio cap and no price (it links to the Azure Speech pricing page). Read on 2026-10-08.
So the number you can quote is the news post's, and it is a vendor claim. Microsoft's post does not say here what the 150 ms includes, so measure first-byte time from your own region.
What each page says
The news post also claims 55% faster model inference and roughly 60% lower cost than comparable models. The Learn page says MAI-Voice-2.1 (the non-Flash model) prioritizes naturalness and expressivity over latency-critical scenarios, which is Microsoft's own guidance to pick Flash for live calls.
| Claim | News post | Learn page |
|---|---|---|
| Price | $15 per 1M characters | Not stated (links to pricing page) |
| End-to-end latency | 150 ms | Not stated |
| Audio per request | 45 s | Not stated |
| Languages | 23 languages, 26 locales | 23 languages |
| Preview or SLA | Not stated | Public preview, no SLA |
| Regions | Not stated | 14 regions |
What 45 seconds means for scripts
A 45-second cap bounds each synthesis call to a few hundred words. At a typical 150 words per minute (an assumption, not a Microsoft figure), 45 seconds is about 112 words or 650 to 700 characters. Longer scripts have to be split by sentence and the audio joined, which is easy for chat turns and awkward for an audiobook; Microsoft points long-form work at MAI-Voice-2.1 instead.
Where Sume differs
Sume's TTS 1.0 is not a streaming route. The API reference describes it as an async job with polling or a webhook, and says mode: sync is a bounded wait of up to 30 seconds on the HTTP request, not a latency promise. A job can be up to 20,000 characters and 1,200 seconds of audio, so the shape is closer to "render this narration" than "answer this caller".
That makes Sume the wrong tool for a phone agent's first word, and the right one for pre-rendered prompts, ad reads and voice-overs. The latency-by-job-type post sorts which jobs need which.
How to measure it yourself
Time three things separately: the delay from request to first audio byte, the delay to the last byte, and the audio length. A model that reports 150 ms to start but streams slower than real time will still stall a call. Run at least 50 requests at the hours your callers actually use, and report the median and the slowest 5 percent, not the best run.
For Sume's async route the equivalent measurement is the time from submit to a completed job with a downloadable file; the benchmark post shows a Python loop for it. The two numbers answer different questions, so do not put them in one column of a vendor comparison.
Sources
Related posts
More in Models
- Nano Banana 2.1 vs Pro from 512 to 4K: price per image on Sume
Nano Banana 2.1 runs $0.075 at 512 to $0.20 at 4K; Nano Banana Pro is $0.1875 up to 2K and $0.375 at 4K. Full tier table and a per-1,000 view.
- Same Omni video twice: no seed, but a replay returns it
Omni Flash lists no temperature or seed, and Sume rejects seed on every video model. To repeat a clip, replay an idempotency key or use a video_url edit.
- Seedance 2.5 has a 4-second minimum: the shortest clip is $1.0747
Seedance 2.5 on Sume takes 4 to 30 seconds. What a 4-second clip costs at 480p, 720p and 1080p on each Seedance tier, and where shorter hooks go.
- sume/music-auto or pinned lyria-3.5: what changes in the job
Music Router takes sume/music-auto (Lyria 3.5 today), lyria-3.5 or lyria-3-pro. What job.model and routed_model echo back, and why the price stays the same.
Written by Sume