MAI-Voice styles via express-as vs Sume's emotion field
MAI voices set styles with SSML mstts:express-as; some only have neutral. Sume has no SSML; generation_config.emotion is a free string up to 64 characters.

On MAI voices, speaking style is a per-voice list that you apply with SSML mstts:express-as, and some voices, such as the en-IN ones on the voices page, only offer neutral. Sume has no SSML field at all; the closest control is generation_config.emotion, an optional free string of 1 to 64 characters, next to numeric speed and volume.
How each side exposes delivery
Microsoft's voices page says styles vary by voice and are set with SSML, so your first step is to look up the style list for the exact voice you chose. A style name that the voice does not have is a configuration error on your side.
Sume's request schema keeps delivery in generation_config: volume from 0.5 to 2, speed from 0.6 to 1.5, and emotion as a short guide. There is also a deprecated speed enum, and pronunciation_dict_id for words you need said a certain way.
| Control | MAI voices | Sume TTS |
|---|---|---|
| Style or emotion | SSML mstts:express-as, per voice list | generation_config.emotion, free string up to 64 characters |
| Speed | Per the SSML or API options | generation_config.speed, 0.6 to 1.5 |
| Volume | Per the SSML or API options | generation_config.volume, 0.5 to 2 |
| Voice with no styles | neutral only, for example some en-IN voices | Emotion string accepted; effect not guaranteed |
| Markup | SSML | None; plain transcript |
What to expect from a free-form emotion
Because the Sume field is a hint rather than an enum, a value that the voice does not respond to simply has little effect. That is more forgiving than a style list and also less predictable, so test each emotion you plan to use on the exact voice. Keep the string short and concrete.
Do not copy SSML into the transcript. Sume has no SSML field, so tags in the text are not interpreted as markup; the API only removes caption markup such as || before the voice reads.
Limits
This comparison covers the control surface, not quality. The MAI models are public preview with no SLA, so style lists may change, while the Sume field limits are from the current request schema. See jobs and results for reading the finished audio.
A migration pattern
If you currently use styles, keep them as data rather than markup, so the target can change without a rewrite.
- Store each line with a plain-language mood such as calm or upbeat, not with the vendor's tag.
- Map that mood to the vendor field at request time: a style name for MAI, a short emotion string for Sume.
- Check
neutral-only voices, which ignore your mood on the MAI side. - Re-listen after any model change; the same word does not guarantee the same delivery.
Sources
Related posts
More in Comparisons
- Max video length in October 2026: TikTok, YouTube Shorts, Sume tools
TikTok's API allows 10 minutes and YouTube Shorts 3. Sume caps timeline and inspect at 1800 s, trim output at 900 s, and frames and filter at 300 s.
- Meta signs the EU AI labeling code: an advertiser checklist
Meta confirmed on July 28, 2026 it will sign the EU AI-content transparency code. No ad deadline is stated, and AI info labels already exist on ads.
- Microsoft's October 2026 speech launches: what to change in a pipeline
MAI-Voice-2.1, its Flash model and MAI-Transcribe-2-Streaming, claim by claim, with what each means for an ad voice pipeline and what Sume covers today.
- Midjourney edits now touch only selected pixels; the API route on Sume
Midjourney's Sep 24 update says edits modify only selected pixels. Midjourney isn't on Sume; a masked edit via mask_url on ChatGPT Image 2.5 is the API route.
Written by Sume