MAI-Voice-2.1 zero-shot voice prompting vs a Sume voice id or avatar
MAI-Voice-2.1 lists zero-shot voice prompting. Sume TTS takes a voice id or an avatar's voice, not a sample clip. What each request can and cannot express.
Microsoft's MAI-Voice-2.1 page lists "Zero-shot Voice Prompting" and "Instant voice matching" with no fine-tuning, for both the standard and Flash models. Sume TTS has no equivalent request field: a TTS request selects a voice by voice.id, or by avatar_id or avatar_handle, and none of the three takes a reference audio clip. If you need a voice matched from a sample, price the MAI route separately; Sume does not list MAI.
What the Microsoft page says
The model page, read 2026-10-09, shows a feature matrix for MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Both rows say yes to granular emotion control and zero-shot voice prompting, and both are described as supporting instant voice matching. The page lists 23 languages. The page gives no sample length, consent flow or licensing terms in the text read for this post, so those are open questions to settle with Microsoft before you clone anyone's voice.
What a Sume TTS request can select
The request schema needs exactly one text source (transcript or transcript_source) and a voice selector. The voice id must be a UUID (8-4-4-4-12 hex) or a Voices library id (voi_ plus 32 hex), not a name such as alloy. An avatar can supply the voice instead: list GET /v1/avatar-1.0/avatars and use an avatar whose voice.status is ready. Provider credential fields such as api_key are refused, because the surface authenticates with your Sume key only.
| Need | MAI-Voice-2.1 (Microsoft page) | Sume TTS 1.0 |
|---|---|---|
| Match a voice from a short sample | Zero-shot voice prompting: yes | No field for a sample clip |
| Pick a stock voice | Not stated on the page text read | voice.id as UUID or voi_ id |
| Reuse an avatar's voice | Not applicable | avatar_id or avatar_handle (voice.status ready) |
| Reproduce a voice on the next line | Not stated | Read voice, language, speed and generation_config from the finished job |
Keeping a voice consistent without cloning
A completed Sume text-to-speech job records how the audio was made: model_id, voice as { "mode": "id", "id": ... }, language, output_format, generation_config and speed, each null when the request did not send it. The jobs docs say to read these values from the job so the next line sounds the same. For a series, store that record with the episode and reuse it.
For a narrator that is tied to a character, use the avatar's handle, so the voice and the face stay together across talking video and TTS. The cost is the same per-character rate: $0.0475 per 1,000 characters, so a 900-character line is $0.04275.
A decision rule
Pick by the thing you must reproduce. Whichever engine you use, a voice that sounds like a real person is a consent and disclosure question before it is a technical one; keep the permission record next to the job id, and do not use a stock voice to imply that a named person said the line.
- Need a voice from an arbitrary sample, now: use the engine that lists it, and keep the consent record.
- Need a repeatable narrator for a channel: pin one voice id and store the job's recorded settings.
- Need a presenter who also appears on screen: use an avatar and its handle.
Sources
Related posts
More in Comparisons
- MiniMax H3 768p vs Gemini Omni Flash 720p: an 8-second clip
On Sume, an 8-second clip is $0.60 on minimax-h3 at 768p and $1.00 on gemini-omni-flash-1.1 at 720p, with $0.30 at Omni's 360p. Price and constraints.
- MiniMax H3 Max vs Wan 3.0 at 1080p: 5, 10 and 15 seconds on Sume
On Sume, MiniMax H3 Max at 1080p costs $0.20 a second and Wan 3.0 costs $0.25. For 5, 10 and 15 seconds: $1.00/$2.00/$3.00 vs $1.25/$2.50/$3.75.
- MiniMax H3 Max 768p: 15 s is $1.50 on Sume, $1.20 on Runway Dev
Runway Dev lists h3_max at 5 credits per second at 480p and 8 at 768p. Sume lists minimax-h3-max at $0.0625 and $0.10, plus a 1080p row at $0.20.
- MiniMax Hailuo 2.3 at 56 cents vs H3 at 80 cents, 10 seconds of 768P
MiniMax's page lists Hailuo 2.3 under Legacy Models: $0.56 for a 10-second 768P clip, $0.32 for Fast. H3 is $0.80. Sume lists H3 only.
Written by Sume