MAI-Voice-2.1 emotion control vs Sume TTS emotion, speed and volume

MAI-Voice-2.1 lists emotion control. Sume TTS takes generation_config: emotion (1-64 chars), speed 0.6-1.5, volume 0.5-2. A three-take test costs 3 cents.

5 min readSume
All posts

Microsoft lists "Granular Emotion Control" for MAI-Voice-2.1 and its Flash version. On Sume TTS, delivery is steered by an optional generation_config object: emotion (a string of 1 to 64 characters), speed (0.6 to 1.5) and volume (0.5 to 2), plus a separate top-level speed of slow, normal or fast. A three-take emotion test of a 200-character line costs about 3 x $0.0095 = $0.0285.

What each side documents

The Microsoft page, read 2026-10-09, states the feature as a yes for both models. The text read for this post does not list the emotions, their intensity scale, or how a script marks them, so do not assume a tag syntax carries over from one engine to another.

Sume's schema is explicit about types and ranges, and silent about the vocabulary. emotion is a free string with a 64-character limit; which values change the audio is a property of the engine behind the request. The only reliable way to learn the vocabulary is to try a short line and listen.

Delivery controls, Sume request schema and Microsoft page, read 2026-10-09
ControlSume TTS 1.0MAI-Voice-2.1
Emotiongeneration_config.emotion, string 1-64 charactersGranular emotion control: yes (vocabulary not in the page text read)
Speaking rategeneration_config.speed 0.6-1.5, or speed slow / normal / fastNot stated in the page text read
Loudnessgeneration_config.volume 0.5-2Not stated in the page text read
What was usedEchoed on the finished job; null when not sentNot applicable

A cheap three-take test

Pick one 200-character line and run it three times with the same voice and language, changing only emotion. Sume charges $0.0475 per 1,000 characters, so each take is 200 x 0.0475 / 1,000 = $0.0095, and three takes are $0.0285. Use a different Idempotency-Key for each take, because the same key returns the first job.

Compare on the same speakers and the same room. Then read generation_config back from each finished job and keep the one you chose with the episode, so the next line can reuse it.

Language and voice come first

Emotion is the last knob, not the first. A voice that does not speak the line's language will not be rescued by an emotion string, and Sume's language field (2 to 16 characters) is checked against the voice: a mismatch is refused until you confirm it with confirm_language_mismatch. Set the language and voice, listen once with no emotion at all, and only then add generation_config to move from a flat read to the one you want. Store what you used; the finished job echoes it back.

Mistakes to avoid

Four habits keep an emotion test honest.

  • Do not send both a top-level speed and generation_config.speed in the same request until you have tested what you get; pick one control.
  • Do not put emotion words into the transcript and expect them to be performed. The transcript is spoken text.
  • A 64-character limit means a label, not a sentence. Keep it short.
  • Do not judge a voice on one line. Emotion that suits a 12-word hook can sound strained over 400 characters.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume