MAI-Voice-2 voice prompting from 5-60 s of audio vs Sume TTS
Microsoft MAI-Voice-2 prompts a voice from 5 to 60 seconds of reference audio. Sume TTS selects an existing voice and takes no reference clip.

What is MAI-Voice-2 voice prompting, and does Sume have it?
Voice prompting in MAI-Voice-2 means giving the model a short reference clip and getting speech in that voice back, with no retraining. Microsoft says it works with 5 to 60 seconds of reference audio in all supported languages. Sume TTS has no equivalent: the request takes an avatar or a voice id, and the OpenAPI contract documents no reference-audio field.
The Microsoft AI announcement is dated June 2, 2026, and the model is generally available in Microsoft Foundry. This post reads that page, not later MAI releases.
What does Microsoft say MAI-Voice-2 can do?
The page lists 15 languages and locales, among them English (US, Australia), Spanish (Spain, Mexico), Portuguese (Brazil, Portugal), Korean, Hindi and Simplified Chinese. It describes emotion control through tags such as sad, whispered and excited, code-switching for Hindi-English and Spanish-English pairs, and a stable speaker identity across long-form content like audiobooks and lectures.
It also states a safety rule: "Only authorized, licensed voices can be synthesized in production. No unlicensed voice cloning is possible."
- 5 to 60 seconds of reference audio for voice prompting.
- Custom voices created in Foundry from short clips, without retraining.
- Emotion tags and code-switching between listed language pairs.
- Production use limited to authorized, licensed voices.
How does voice choice work on Sume instead?
A Sume TTS request picks the voice with avatar_id, avatar_handle or voice.id. The id must be a TTS voice UUID or a Voices library id starting voi_. A reference clip, a consent recording or a prompt audio URL is not part of the contract, and the contract tells callers to authenticate with the Sume API key only.
| Question | MAI-Voice-2 | Sume TTS 1.0 |
|---|---|---|
| Reference audio | 5 to 60 seconds | Not accepted |
| Languages | 15 languages and locales listed | Set per request with language |
| Emotion | Tags such as sad, whispered, excited | generation_config.emotion, free text |
| Who may be synthesized | Authorized, licensed voices only | Existing workspace voices only |
| Where it runs | Microsoft Foundry | api.sume.com job API |
What should you do if you need a voice from a clip?
Create it where the vendor verifies authorization, render the narration there, and bring the file into Sume for the parts Sume does well: joining takes with timeline audio, burning captions, and assembling the final video. Sume joins hosted audio without re-synthesising it, so the voice stays exactly as exported.
If a stock voice is acceptable, stay on Sume end to end and fix one avatar handle per project. The voice cloning post explains why a stable voice identity matters more for ads than a clone.
What is the checklist before you ship?
Keep the consent paperwork for any cloned voice with the project. Confirm the language of each script, since Sume needs it set explicitly for non-English text. Listen to a 20-second sample from each system before comparing price, because per-character prices only matter once the voice is acceptable.
Sources
Related posts
More in Comparisons
- Advantage+ background generation for catalog ads: Meta's 2-3% claim
Meta says background generation for catalog ads lifted conversions 2-3%. You can turn it off; to control the scene yourself, edit the product shot on Sume.
- Meta Muse for small business: approval before spend vs Sume MCP
Meta's Muse for small business asks approval before publishing, sending or spending. Sume's MCP uses an idempotency key, which is not human approval.
- Meta One creator plans: $14.99 to $499 a month, what is in them
Meta One launched September 15, 2026 with creator plans from $14.99 to $499 a month. What Reels creators get, and what Sume does and does not replace.
- MiniMax Speech 2.8 pitch and emotions vs Sume TTS controls
MiniMax T2A lists speech-2.8 models, nine emotions, pitch -12 to 12 and 10,000 characters. Sume TTS has speed, volume and emotion text, no pitch.
Written by Sume