MAI-Voice styles via express-as vs Sume's emotion field

MAI voices set styles with SSML mstts:express-as; some only have neutral. Sume has no SSML; generation_config.emotion is a free string up to 64 characters.

4 min readSume
All posts

On MAI voices, speaking style is a per-voice list that you apply with SSML mstts:express-as, and some voices, such as the en-IN ones on the voices page, only offer neutral. Sume has no SSML field at all; the closest control is generation_config.emotion, an optional free string of 1 to 64 characters, next to numeric speed and volume.

How each side exposes delivery

Microsoft's voices page says styles vary by voice and are set with SSML, so your first step is to look up the style list for the exact voice you chose. A style name that the voice does not have is a configuration error on your side.

Sume's request schema keeps delivery in generation_config: volume from 0.5 to 2, speed from 0.6 to 1.5, and emotion as a short guide. There is also a deprecated speed enum, and pronunciation_dict_id for words you need said a certain way.

Delivery controls (Microsoft voices page read 2026-10-05; Sume from the repo)
ControlMAI voicesSume TTS
Style or emotionSSML mstts:express-as, per voice listgeneration_config.emotion, free string up to 64 characters
SpeedPer the SSML or API optionsgeneration_config.speed, 0.6 to 1.5
VolumePer the SSML or API optionsgeneration_config.volume, 0.5 to 2
Voice with no stylesneutral only, for example some en-IN voicesEmotion string accepted; effect not guaranteed
MarkupSSMLNone; plain transcript

What to expect from a free-form emotion

Because the Sume field is a hint rather than an enum, a value that the voice does not respond to simply has little effect. That is more forgiving than a style list and also less predictable, so test each emotion you plan to use on the exact voice. Keep the string short and concrete.

Do not copy SSML into the transcript. Sume has no SSML field, so tags in the text are not interpreted as markup; the API only removes caption markup such as || before the voice reads.

Limits

This comparison covers the control surface, not quality. The MAI models are public preview with no SLA, so style lists may change, while the Sume field limits are from the current request schema. See jobs and results for reading the finished audio.

A migration pattern

If you currently use styles, keep them as data rather than markup, so the target can change without a rewrite.

  • Store each line with a plain-language mood such as calm or upbeat, not with the vendor's tag.
  • Map that mood to the vendor field at request time: a style name for MAI, a short emotion string for Sume.
  • Check neutral-only voices, which ignore your mood on the MAI side.
  • Re-listen after any model change; the same word does not guarantee the same delivery.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume