MAI-Voice-2.1 or Flash for explainer narration: latency barely matters

Microsoft lists about 550 ms vs 45 ms model inference. For narration rendered ahead of time, pick on quality and price, then see how Sume jobs return audio.

5 min readSume
All posts

For narration you render before anyone is listening, latency is not the deciding factor. Microsoft's MAI-Voice-2.1 page lists about 550 ms of model inference for the standard model and about 45 ms for Flash, and points the standard one at audiobooks, content creation and voice-over where fidelity comes first, and Flash at call centers, assistants and IVR. So for an explainer video choose the standard model unless your cost target forces Flash, then judge both by ear.

The differences Microsoft states

Both variants support 23 languages and emotion control. The page differs on speed, price and intended use.

MAI-Voice-2.1 and MAI-Voice-2.1-Flash as stated by Microsoft (read 2026-10-04)
ItemMAI-Voice-2.1MAI-Voice-2.1-Flash
Model inferenceAbout 550 msAbout 45 ms
Price$22 per million characters$15 per million characters
Positioned forAudiobooks, content, voice-overCall centers, assistants, IVR
Languages23 languagesSame

What the price gap is worth

For a 5,000-character script, about five minutes of speech, the standard model is $0.11 and Flash is $0.075 at list price. A difference of 3.5 cents per script rarely beats a better take. If you voice thousands of scripts a month, the gap becomes a budget line, and an A/B listen on your own scripts is the only honest way to decide.

Where Sume comes in

The Sume TTS route is an async job that returns a hosted audio file, so latency to first sound is not the metric; total job time and cost are. Sume TTS 1.0 is billed per transcript character at Cartesia's list rate times 1.25, which works out to $0.0475 per 1,000 characters, or $0.2375 for the same 5,000 characters. Microsoft's MAI-Voice models are not part of that: the TTS router's documented engines are Cartesia Sonic ids, so read GET /v1/tts-router/models rather than assuming, as the API reference describes.

Once you have a narration file, the rest of a video is assembly: cut scenes to the voice with Timeline 1.0, or join takes with timeline audio.

A short audition

Voice the same 400-character paragraph with both variants. Compare the opening word, how each reads a number, and whether the tone holds through the last sentence. Pick the cheaper one only if you cannot hear the difference. See jobs and results for reading finished jobs.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume