MAI-Voice-2.1 ~550 ms vs Flash ~45 ms: does it matter for voiceover?
Microsoft lists ~550 ms and ~45 ms model-inference latency for MAI-Voice-2.1 and Flash. In series, 1,000 lines take 9.2 minutes vs 45 s. Read 2026-10-08.

Microsoft's model page lists MAI-Voice-2.1 at about 550 ms of model-inference latency and MAI-Voice-2.1-Flash at about 45 ms, and it labels the first for fidelity (audiobooks, content creation, voice-over) and the second for latency-sensitive uses (call-center agents, voice assistants, IVR). For a voiceover that someone listens to later, the difference is invisible: it is the delay before the first audio, not the length of the audio. It matters only when a person or a phone line is waiting.
What the page numbers measure
The page labels both figures "Latency (model inference)". That is a model-side figure; it excludes your network, queueing and any download of the finished file, and the page does not give a figure for a long script. Treat the numbers as a ranking between the two models rather than a promise about your request.
A queue of voiceover lines is the case where inference time adds up. If the lines were processed strictly one after another, 1,000 lines would spend 1,000 x 0.55 s = 550 s (9.2 minutes) in the 2.1 model and 1,000 x 0.045 s = 45 s in Flash, before any network time. That is an illustration of the page's numbers, not a measured throughput; parallel requests change it.
| Model | Latency (model inference) | Price | Page's best-for list | 1,000 lines in series (inference only) |
|---|---|---|---|---|
| MAI-Voice-2.1 | ~550 ms | $22 per 1M chars | Fidelity matters more than speed: audiobooks, content creation, voice-over | 9.2 min |
| MAI-Voice-2.1-Flash | ~45 ms | $15 per 1M chars | Latency sensitive: call-center agents, voice assistants, IVR | 45 s |
How a Sume TTS job is shaped
Sume TTS is a job, not a streaming call. The default mode is async: the request returns 202 with a job envelope and you poll its status until it is terminal, then read the result. A sync or subscribe mode waits for up to 30 seconds on the submit call and, if the job has not finished, you keep polling rather than resubmitting (Jobs and results). The catalog prices a TTS request at $0.0475 per 1,000 characters with a 20,000-character maximum.
That shape fits batch voiceover, where you queue lines and collect results, and it does not fit a phone agent that must start speaking within a fraction of a second. Sume does not list a MAI model, so if you need the Flash latency profile the choice is a separate provider, not a Sume setting.
A rule for choosing
Ask what is waiting. If a person hears the first syllable as soon as the model emits it, latency decides and the price gap per million characters ($22 versus $15) is secondary. If the audio goes into a video, the full-length result is what you need, a 550 ms start is irrelevant and the question becomes price, voice and how the result joins the edit.
Three checks before you pick on latency
Latency is one of four things that separate the two models on Microsoft's page; the others are price ($22 versus $15 per million characters), the best-for list, and fidelity, which the page ties to the slower model. If your audio is for a video, the useful comparison is therefore cost per finished minute at the fidelity you can hear, and a blind listen on your own script settles it faster than a latency figure.
- Is anyone listening while the model works? If not, latency is a throughput question, not a quality one.
- Is your queue serial? Parallel requests hide most of the per-request delay, so the 9.2-minute figure is a worst case for in-series work.
- Does the result have to be stored? A Sume job returns a media.sume.com artifact you can hand to the next step, which matters more for a pipeline than the first-byte time.
Sources
Related posts
More in Models
- MAI-Voice-2.1-Flash: 150 ms end-to-end and a 45-second audio limit
Microsoft's news post gives MAI-Voice-2.1-Flash 150 ms end-to-end latency and 45 s of audio; the Learn page gives no latency number. What each page supports.
- MiniMax H3 4K hero clip vs 768p cutdowns: holiday ad cost
A 10-second MiniMax H3 clip is $0.75 at 768p, $1.625 at 2K and $2.00 at 4K on Sume. Five free reference images are included; extra ones are $0.10 each.
- MiniMax H3 defaults to 2K at MiniMax; on Sume the default is 768p
MiniMax's H3 page says 2K is the default. Sume's minimax-h3 runs native 480p or 768p with a 768p default; 2K and 4K are billed upscales.
- MiniMax H3 on a 12 GB RTX 3060: 480p only, 42 GB, and the price
MiniMax says 12 GB of VRAM runs 480p H3 clips with audio from a pruned int8 checkpoint of about 42 GB. What that buys, and a 480p clip on Sume costs.
Written by Sume