MAI-Voice-2.1 zero-shot voice prompting vs a Sume voice id or avatar

MAI-Voice-2.1 lists zero-shot voice prompting. Sume TTS takes a voice id or an avatar's voice, not a sample clip. What each request can and cannot express.

5 min readSume
All posts

Microsoft's MAI-Voice-2.1 page lists "Zero-shot Voice Prompting" and "Instant voice matching" with no fine-tuning, for both the standard and Flash models. Sume TTS has no equivalent request field: a TTS request selects a voice by voice.id, or by avatar_id or avatar_handle, and none of the three takes a reference audio clip. If you need a voice matched from a sample, price the MAI route separately; Sume does not list MAI.

What the Microsoft page says

The model page, read 2026-10-09, shows a feature matrix for MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Both rows say yes to granular emotion control and zero-shot voice prompting, and both are described as supporting instant voice matching. The page lists 23 languages. The page gives no sample length, consent flow or licensing terms in the text read for this post, so those are open questions to settle with Microsoft before you clone anyone's voice.

What a Sume TTS request can select

The request schema needs exactly one text source (transcript or transcript_source) and a voice selector. The voice id must be a UUID (8-4-4-4-12 hex) or a Voices library id (voi_ plus 32 hex), not a name such as alloy. An avatar can supply the voice instead: list GET /v1/avatar-1.0/avatars and use an avatar whose voice.status is ready. Provider credential fields such as api_key are refused, because the surface authenticates with your Sume key only.

Voice selection, MAI page and Sume request schema, read 2026-10-09
NeedMAI-Voice-2.1 (Microsoft page)Sume TTS 1.0
Match a voice from a short sampleZero-shot voice prompting: yesNo field for a sample clip
Pick a stock voiceNot stated on the page text readvoice.id as UUID or voi_ id
Reuse an avatar's voiceNot applicableavatar_id or avatar_handle (voice.status ready)
Reproduce a voice on the next lineNot statedRead voice, language, speed and generation_config from the finished job

Keeping a voice consistent without cloning

A completed Sume text-to-speech job records how the audio was made: model_id, voice as { "mode": "id", "id": ... }, language, output_format, generation_config and speed, each null when the request did not send it. The jobs docs say to read these values from the job so the next line sounds the same. For a series, store that record with the episode and reuse it.

For a narrator that is tied to a character, use the avatar's handle, so the voice and the face stay together across talking video and TTS. The cost is the same per-character rate: $0.0475 per 1,000 characters, so a 900-character line is $0.04275.

A decision rule

Pick by the thing you must reproduce. Whichever engine you use, a voice that sounds like a real person is a consent and disclosure question before it is a technical one; keep the permission record next to the job id, and do not use a stock voice to imply that a named person said the line.

  • Need a voice from an arbitrary sample, now: use the engine that lists it, and keep the consent record.
  • Need a repeatable narrator for a channel: pin one voice id and store the job's recorded settings.
  • Need a presenter who also appears on screen: use an avatar and its handle.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume