MAI-Voice-2.1 whispering and shouting: the 19 voices that list both
Only 19 MAI-Voice-2.1 voices list whispering and shouting by our count. English-UK and Korean voices do not. Sume has no whisper switch.

By our count of Microsoft's MAI-Voice-2.1 page, 19 voices list both whispering and shouting, including en-US Ethan, Harper and Olivia. The two en-GB voices and the Korean voices do not. Sume TTS has no whisper or shout setting; you can only hint through the emotion text and volume.
Voices that list both
Counted from the Learn page:
| Locale | Voices |
|---|---|
| de | Klaus, Mia |
| en-AU | Isla |
| en-US | Ethan, Harper, Olivia |
| es-MX | Alejo, Valeria |
| fr | Marc, Soleil |
| hi | Dhruv, Kavya, Priya |
| it | Luca, Rosa |
| pt-BR | Caio, Luana |
| zh-CN | Bo, Mei |
Gaps to plan around
en-GB Emily and Harry do not list the pair, and the Korean voices list softvoice only. A whispered ASMR-style product spot in British English therefore has no listed option in this family. Re-check the page before locking a cast, since this is a preview.
What Sume can do
Sume's generation_config has volume from 0.5 to 2.0 and a 64-character emotion string (API reference). Writing whispering, close to the mic may nudge the delivery, but we have not verified how the model responds, so test a short line and keep the clip only if you like it. Lowering volume makes it quieter, not whispered.
Sources
Related posts
More in Models
- MAI-Voice 50.3% of 4,000 listeners: how to quote the Turing claim
Microsoft said 50.3% of 4,000 listeners rated MAI-Voice as equally or more human-like than human recordings. What it covers, and a cheap test.
- Three characters, three dances: Omni IMAGE_REF and VIDEO_REF tokens
Google's Omni 1.1 demo swaps three dancers for a dog, an octopus and a bear. Here is the same request on Sume, with the 0-based reference tokens in order.
- MiniMax H3 outputs 24 fps and 32 kHz stereo: what Sume's H3 adds
MiniMax's H3 open-source post lists 24 fps video and 32 kHz stereo audio, 4 to 15 s clips and 768p by default. Sume's minimax-h3 takes 5 to 15 s.
- MiniMax H3 or H3 Max after Sora: native 768p vs latent 1080p prices
On Sume, H3 renders native 480p or 768p and bills 2K/4K upscales; H3 Max adds 1080p as a latent refinement of 768p. 10-second prices side by side.
Written by Sume