Which MAI-Voice-2.1 model for ad voiceover: standard or Flash?
Microsoft positions MAI-Voice-2.1 for voice-over and audiobooks, Flash for live agents. For rendered ads pick standard; Sume is the async file path.

For ad voiceover pick standard MAI-Voice-2.1, not Flash. Microsoft's own model page puts audiobooks, content creation and voice-over with the standard model, and call-center agents, voice assistants and IVR with Flash.
Microsoft's split
Microsoft's pages describe Flash as the low-latency variant: about 45 ms model latency at $15 per million characters. Standard is listed at about 550 ms and $22 per million.
| Property | MAI-Voice-2.1 | MAI-Voice-2.1-Flash |
|---|---|---|
| Price per 1M characters | $22 | $15 |
| Latency on model page | About 550 ms | About 45 ms |
| Microsoft's suggested use | Audiobooks, content creation, voice-over | Call-center agents, voice assistants, IVR |
| Languages | 23 | 23 |
What an ad pipeline needs
An ad is a script, a voice and a finished file. Nobody waits for the first syllable, so latency does not matter. What matters is that the voice can be matched to a brand, takes can be repeated and the file lands in an editing timeline.
The Sume path for finished files
Sume's TTS Router is built around async jobs: POST /v1/tts-router/generate returns a request id, you poll status and read the result, and a job covers up to 1,200 seconds of audio. There is no streaming and no SSML field. Output defaults to mp3 at 44.1 kHz and 128 kbps; output_format also takes wav or raw.
curl -sS -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"sonic-3.6","transcript":"Fresh bread, baked at six.","avatar_handle":"product_host","language":"en","output_format":{"container":"wav","sample_rate":48000,"encoding":"pcm_s16le"}}' \
| jq '.data | {request_id, status_url, result_url}'When Flash is the right call
If the same script is read live to a caller, Flash fits and Sume does not: Sume has no streaming. Call-center and IVR work is Microsoft's own listed use for it.
Sources
Related posts
More in Comparisons
- Text change, region change or background swap: which Sume image route
Pick the Sume image edit route by the kind of change: Ideogram 4.5 for words, ChatGPT Image 2.5 with mask_url for a region, a reference edit for backgrounds.
- Which platforms auto-detect AI media: YouTube, Meta, Pinterest, TikTok
YouTube, Meta and Pinterest say they can label AI media from metadata or detection; TikTok auto-labels its own AI effects. None replaces your own disclosure.
- Who owns a Suno song? What 'Output owned by Suno' means for ads
Suno's terms assign Pro and Premier users the Output Suno owns, but make no copyright warranty. What to file with an ad, and how a Sume music job record helps.
- Wondercraft Creator $25 vs a podcast intro built on Sume for $0.24
Wondercraft lists Creator at $25 and Pro at $45. A 60-second show intro from text to speech, generated music and one render costs about $0.24 on Sume.
Written by Sume