MAI-Voice-2.1 SSML style="happiness" is not in the voice style list
Microsoft's MAI-Voice-2.1 SSML example uses style="happiness", but the Learn page's own style lists say happy or joyful. Check styles before you ship.

If MAI-Voice-2.1 ignores or rejects <mstts:express-as style="happiness">, check the style name before debugging anything else. The Learn page's expressive-control example uses style="happiness" on en-US-Harper, yet the same page's managed-voice table lists Harper's styles as agent, angry, audiobook, confused, customer_call_center, determined, educational, embarrassed, excited, happy, hopeful, joyful, narrator, neutral, regretful, relieved, sad, shouting, softvoice and whispering. happiness is not among them (read 2026-10-08).
I did not run Microsoft's service. The point is narrower: the documentation contradicts itself, so copy style names from the table, not from the example.
What the page shows
The page states that developers control expression through mstts:express-as and style, with emotions such as joy, excitement and empathy. Those three words also differ from the table names (joyful, excited, caringempathy). Three places, three spellings.
Style availability also varies by voice. Some voices have only neutral; others list 19 styles; some Spanish (Spain), Dutch, Russian, Thai and Turkish voices in the table use a different set (adventurous, caringempathy, curious and so on).
| Where the page uses it | Style string | In Harper's table row |
|---|---|---|
| Example SSML | happiness | No |
| Feature description | joy | No (joyful is) |
| Feature description | excitement | No (excited is) |
| Feature description | empathy | No |
| Table | happy | Yes |
| Table | joyful | Yes |
A safe way to test
Build the SSML from the table value, synthesize one short sentence, and listen for a difference against neutral. Keep the sentence identical across styles so any change is the style. Treat the model as a public preview: the page says the preview has no service-level agreement, so re-run the test when you upgrade the SDK.
<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis"
xmlns:mstts="http://www.w3.org/2001/mstts" xml:lang="en-US">
<voice name="en-US-Harper:MAI-Voice-2.1">
<mstts:express-as style="happy">
Your order has shipped.
</mstts:express-as>
</voice>
</speak>What Sume exposes instead
Sume's POST /v1/tts-1.0/generate has no SSML. Tone goes through generation_config: volume from 0.5 to 2.0, speed from 0.6 to 1.5, and emotion, a free-text guide of up to 64 characters. Because emotion is free text, there is no list of names to get wrong; the trade is that you cannot rely on a fixed vocabulary either.
Per-voice style enumeration like Microsoft's table is not part of Sume's public contract, so test each voice by ear, as above. For a 240-character line, a retake costs two cents at $0.0475 per 1,000 characters.
A note on documentation drift
Preview documentation changes quickly. The Learn page carries an update date of 2026-10-01 in its metadata for the voices page, and a style rename or an added style between now and general availability would not surprise anyone. Keep the style name in one constant in your code, not scattered through SSML strings, so a rename is a one-line change.
If a style is silently ignored, the output is simply the neutral voice, which can be hard to notice in a long script. Add a test that synthesizes the same sentence in neutral and in your chosen style and fails if the two files are byte-identical. It costs a fraction of a cent per run and turns a quiet doc mismatch into a loud one.
Sources
Related posts
More in Developers
- Modal 1.6.1 endpoint logs and stats: debug a Sume webhook receiver
Modal 1.6.1 adds modal endpoint info, stats and logs. Use them to see why a Sume job webhook got a 401 or a timeout, and check the 150 s web timeout first.
- model sume/auto on /v1/videos: defaults, limits, replay-stable price
sume/auto lets Sume pick the video family. Defaults are 720p and 8 s, clips run 3 to 10 s at 16:9 or 9:16, and a replay gets the same route and price.
- Music API 400s: duration and negative_prompt, and the fix
Sume's Music Router returns 400 if you send duration, duration_seconds or a non-empty negative_prompt. The fix for each, and a request that passes validation.
- Music prompt limits: 1-5000 characters and no duration field
Sume's Music Router takes a 1 to 5000 character prompt and rejects duration. How to set length and sections in the prompt text, with a request.
Written by Sume