MAI-Voice-2.1 ~550 ms vs Flash ~45 ms: does it matter for voiceover?

Microsoft lists ~550 ms and ~45 ms model-inference latency for MAI-Voice-2.1 and Flash. In series, 1,000 lines take 9.2 minutes vs 45 s. Read 2026-10-08.

5 min readSume
All posts

Microsoft's model page lists MAI-Voice-2.1 at about 550 ms of model-inference latency and MAI-Voice-2.1-Flash at about 45 ms, and it labels the first for fidelity (audiobooks, content creation, voice-over) and the second for latency-sensitive uses (call-center agents, voice assistants, IVR). For a voiceover that someone listens to later, the difference is invisible: it is the delay before the first audio, not the length of the audio. It matters only when a person or a phone line is waiting.

What the page numbers measure

The page labels both figures "Latency (model inference)". That is a model-side figure; it excludes your network, queueing and any download of the finished file, and the page does not give a figure for a long script. Treat the numbers as a ranking between the two models rather than a promise about your request.

A queue of voiceover lines is the case where inference time adds up. If the lines were processed strictly one after another, 1,000 lines would spend 1,000 x 0.55 s = 550 s (9.2 minutes) in the 2.1 model and 1,000 x 0.045 s = 45 s in Flash, before any network time. That is an illustration of the page's numbers, not a measured throughput; parallel requests change it.

Microsoft's listed latency, price and intended use, read 2026-10-08
ModelLatency (model inference)PricePage's best-for list1,000 lines in series (inference only)
MAI-Voice-2.1~550 ms$22 per 1M charsFidelity matters more than speed: audiobooks, content creation, voice-over9.2 min
MAI-Voice-2.1-Flash~45 ms$15 per 1M charsLatency sensitive: call-center agents, voice assistants, IVR45 s

How a Sume TTS job is shaped

Sume TTS is a job, not a streaming call. The default mode is async: the request returns 202 with a job envelope and you poll its status until it is terminal, then read the result. A sync or subscribe mode waits for up to 30 seconds on the submit call and, if the job has not finished, you keep polling rather than resubmitting (Jobs and results). The catalog prices a TTS request at $0.0475 per 1,000 characters with a 20,000-character maximum.

That shape fits batch voiceover, where you queue lines and collect results, and it does not fit a phone agent that must start speaking within a fraction of a second. Sume does not list a MAI model, so if you need the Flash latency profile the choice is a separate provider, not a Sume setting.

A rule for choosing

Ask what is waiting. If a person hears the first syllable as soon as the model emits it, latency decides and the price gap per million characters ($22 versus $15) is secondary. If the audio goes into a video, the full-length result is what you need, a 550 ms start is irrelevant and the question becomes price, voice and how the result joins the edit.

Three checks before you pick on latency

Latency is one of four things that separate the two models on Microsoft's page; the others are price ($22 versus $15 per million characters), the best-for list, and fidelity, which the page ties to the slower model. If your audio is for a video, the useful comparison is therefore cost per finished minute at the fidelity you can hear, and a blind listen on your own script settles it faster than a latency figure.

  • Is anyone listening while the model works? If not, latency is a throughput question, not a quality one.
  • Is your queue serial? Parallel requests hide most of the per-request delay, so the 9.2-minute figure is a worst case for in-series work.
  • Does the result have to be stored? A Sume job returns a media.sume.com artifact you can hand to the next step, which matters more for a pipeline than the first-byte time.

Sources

Related posts

More in Models

All Models posts

Written by Sume