MAI-Voice-2.1-Flash at 45 ms: does a rendered avatar need fast TTS?

Microsoft lists MAI-Voice-2.1-Flash at about 45 ms of inference. A rendered avatar clip does not benefit from it. Where the latency shows up in a Sume job.

4 min readSume
All posts

No. A rendered avatar clip does not need a 45 ms text-to-speech model, because the voice is one stage inside a job that renders video over many seconds, and the viewer never waits on the speech alone. Low latency pays off only when a person is talking to the avatar live.

The numbers Microsoft publishes

Microsoft's MAI-Voice-2.1 page lists about 550 ms of model inference for MAI-Voice-2.1 and about 45 ms for MAI-Voice-2.1-Flash. It describes Flash as the faster, lower-cost variant for call centers, voice assistants and IVR, and the standard model as the fidelity pick for audiobooks and content. Prices are $22 and $15 per million characters.

These are inference figures for the model, not end-to-end times for a product. Microsoft does not claim them for video.

Where time goes in an avatar job

An Avatar 1.0 talking video is an asynchronous job. You POST the script, receive a job id, and then poll, subscribe or take a webhook. The documented window for a request is 4 to 60 seconds of estimated video, and sync mode waits at most 30 seconds before it hands back a queued job. Speech is generated inside that job, so a 500 ms difference in the voice model is lost inside render time that is measured in tens of seconds or more.

Queue time can be larger than either. Sume accepts valid paid jobs as queued while the workspace is at its concurrency limit, and only moves them to processing as slots open.

Which latency matters for which product (vendor figures read 2026-10-05)
ProductWhat the viewer waits onDoes a 45 ms voice help?
Live voice assistant or call agentFirst audio after the caller stops talkingYes, this is the use case
Live conversational avatarAudio and video of the reply togetherPartly; the video render is the other half
Rendered avatar clip (Sume Avatar 1.0)Whole job: queue, speech, render, uploadNo; poll or use a webhook
Narration file for a lip-sync clipAudio job, then lip-sync jobNo; both are async jobs

What to optimise instead

  • Use webhook mode so your server is told when the job finishes, instead of holding a connection open.
  • Preview the first frame before a full render, so a wrong pose costs a still and not a clip.
  • Keep scripts inside the 4 to 60 second window, and split longer scripts into jobs.
  • Set an Idempotency-Key on every paid submit that you might retry.

When to buy a fast voice anyway

If your product is a phone line, a kiosk that answers out loud, or a live agent with a face, the voice latency is the product and a Flash-class model is the right tier. Sume does not run live sessions, so that pipeline lives elsewhere. If you only need a spokesperson who says the same approved lines to many viewers, render once and play the file.

Sources

Related posts

More in Models

All Models posts

Written by Sume