MAI-Voice-2 voice prompting from 5-60 s of audio vs Sume TTS

Microsoft MAI-Voice-2 prompts a voice from 5 to 60 seconds of reference audio. Sume TTS selects an existing voice and takes no reference clip.

5 min readSume
All posts

What is MAI-Voice-2 voice prompting, and does Sume have it?

Voice prompting in MAI-Voice-2 means giving the model a short reference clip and getting speech in that voice back, with no retraining. Microsoft says it works with 5 to 60 seconds of reference audio in all supported languages. Sume TTS has no equivalent: the request takes an avatar or a voice id, and the OpenAPI contract documents no reference-audio field.

The Microsoft AI announcement is dated June 2, 2026, and the model is generally available in Microsoft Foundry. This post reads that page, not later MAI releases.

What does Microsoft say MAI-Voice-2 can do?

The page lists 15 languages and locales, among them English (US, Australia), Spanish (Spain, Mexico), Portuguese (Brazil, Portugal), Korean, Hindi and Simplified Chinese. It describes emotion control through tags such as sad, whispered and excited, code-switching for Hindi-English and Spanish-English pairs, and a stable speaker identity across long-form content like audiobooks and lectures.

It also states a safety rule: "Only authorized, licensed voices can be synthesized in production. No unlicensed voice cloning is possible."

  • 5 to 60 seconds of reference audio for voice prompting.
  • Custom voices created in Foundry from short clips, without retraining.
  • Emotion tags and code-switching between listed language pairs.
  • Production use limited to authorized, licensed voices.

How does voice choice work on Sume instead?

A Sume TTS request picks the voice with avatar_id, avatar_handle or voice.id. The id must be a TTS voice UUID or a Voices library id starting voi_. A reference clip, a consent recording or a prompt audio URL is not part of the contract, and the contract tells callers to authenticate with the Sume API key only.

Voice sourcing, Microsoft page and Sume OpenAPI (read 2026-10-02)
QuestionMAI-Voice-2Sume TTS 1.0
Reference audio5 to 60 secondsNot accepted
Languages15 languages and locales listedSet per request with language
EmotionTags such as sad, whispered, excitedgeneration_config.emotion, free text
Who may be synthesizedAuthorized, licensed voices onlyExisting workspace voices only
Where it runsMicrosoft Foundryapi.sume.com job API

What should you do if you need a voice from a clip?

Create it where the vendor verifies authorization, render the narration there, and bring the file into Sume for the parts Sume does well: joining takes with timeline audio, burning captions, and assembling the final video. Sume joins hosted audio without re-synthesising it, so the voice stays exactly as exported.

If a stock voice is acceptable, stay on Sume end to end and fix one avatar handle per project. The voice cloning post explains why a stable voice identity matters more for ads than a clone.

What is the checklist before you ship?

Keep the consent paperwork for any cloned voice with the project. Confirm the language of each script, since Sume needs it set explicitly for non-English text. Listen to a 20-second sample from each system before comparing price, because per-character prices only matter once the voice is acceptable.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume