MAI-Voice-2.1-Flash at 45 ms: does a rendered avatar need fast TTS?
Microsoft lists MAI-Voice-2.1-Flash at about 45 ms of inference. A rendered avatar clip does not benefit from it. Where the latency shows up in a Sume job.
No. A rendered avatar clip does not need a 45 ms text-to-speech model, because the voice is one stage inside a job that renders video over many seconds, and the viewer never waits on the speech alone. Low latency pays off only when a person is talking to the avatar live.
The numbers Microsoft publishes
Microsoft's MAI-Voice-2.1 page lists about 550 ms of model inference for MAI-Voice-2.1 and about 45 ms for MAI-Voice-2.1-Flash. It describes Flash as the faster, lower-cost variant for call centers, voice assistants and IVR, and the standard model as the fidelity pick for audiobooks and content. Prices are $22 and $15 per million characters.
These are inference figures for the model, not end-to-end times for a product. Microsoft does not claim them for video.
Where time goes in an avatar job
An Avatar 1.0 talking video is an asynchronous job. You POST the script, receive a job id, and then poll, subscribe or take a webhook. The documented window for a request is 4 to 60 seconds of estimated video, and sync mode waits at most 30 seconds before it hands back a queued job. Speech is generated inside that job, so a 500 ms difference in the voice model is lost inside render time that is measured in tens of seconds or more.
Queue time can be larger than either. Sume accepts valid paid jobs as queued while the workspace is at its concurrency limit, and only moves them to processing as slots open.
| Product | What the viewer waits on | Does a 45 ms voice help? |
|---|---|---|
| Live voice assistant or call agent | First audio after the caller stops talking | Yes, this is the use case |
| Live conversational avatar | Audio and video of the reply together | Partly; the video render is the other half |
| Rendered avatar clip (Sume Avatar 1.0) | Whole job: queue, speech, render, upload | No; poll or use a webhook |
| Narration file for a lip-sync clip | Audio job, then lip-sync job | No; both are async jobs |
What to optimise instead
- Use webhook mode so your server is told when the job finishes, instead of holding a connection open.
- Preview the first frame before a full render, so a wrong pose costs a still and not a clip.
- Keep scripts inside the 4 to 60 second window, and split longer scripts into jobs.
- Set an Idempotency-Key on every paid submit that you might retry.
When to buy a fast voice anyway
If your product is a phone line, a kiosk that answers out loud, or a live agent with a face, the voice latency is the product and a Flash-class model is the right tier. Sume does not run live sessions, so that pipeline lives elsewhere. If you only need a spokesperson who says the same approved lines to many viewers, render once and play the file.
Sources
Related posts
More in Models
- MAI-Voice-2.1 Hindi: six voices listed, and a Hindi request on Sume
Microsoft's MAI-Voice-2.1 page lists six hi-IN voices with different style sets. Compare that with a Hindi request on Sume TTS and what the language field does.
- Mandarin Chinese text to speech API: MAI-Voice-2.1 zh-CN vs Sume zh
MAI-Voice-2.1 lists Chinese (Simplified) as zh-CN; Sume tags zh and bills by character. Audition Mandarin ad copy for 1 cent and know what to check.
- Which Foundry models add C2PA and a watermark: GPT Image, FLUX.2, MAI
Microsoft's Foundry page lists the image, audio and voice models that get C2PA credentials and watermarks. What that does and does not tell you about Sume.
- MiniMax H3 API: the cheapest Sume job is 5 s at 480p, $0.3125
On Sume the minimax-h3 model takes 5 to 15 seconds at native 480p or 768p, with 2K and 4K upscales billed on request. Price table by length inside.
Written by Sume