MAI-Voice-2.1 emotion control vs Sume's emotion field

Microsoft lists emotion control on MAI-Voice-2.1. Sume's TTS takes a free-text emotion string, speed and volume. What each gives you, and how to test it.

5 min readSume
All posts

Does Sume's text to speech have the emotion control Microsoft announced for MAI-Voice-2.1? It has an emotion control, but not the same one and not that model. Sume's TTS request takes generation_config.emotion, a free-text guide of up to 64 characters, plus a speed multiplier and a volume multiplier. Microsoft's model is not in Sume's TTS Router, whose catalog is Sonic only.

Microsoft's MAI-Voice-2.1 page (read 2026-10-04) says both MAI-Voice-2.1 and MAI-Voice-2.1-Flash include granular emotion control and zero-shot voice prompting. It does not describe the control surface in the part of the page we could read, so this post compares only what each side states.

What each side documents

The Sume column below comes from the TTS 1.0 request contract in the API reference and its OpenAPI snapshot.

Emotion and delivery controls, Microsoft page and Sume contract, read 2026-10-04
ControlMAI-Voice-2.1 (Microsoft page)Sume TTS 1.0 / Router
EmotionGranular emotion control, no parameter list on the pagegeneration_config.emotion, string, 1-64 characters, described as an emotion guide
SpeedNot stated on the pagegeneration_config.speed, 0.6 to 1.5
VolumeNot stated on the pagegeneration_config.volume, 0.5 to 2.0
Voice from a short clipInstant voice matching from a reference clipVoice chosen by voice.id or an avatar reference; no clone endpoint in this contract

Try an emotion string

The contract does not enumerate accepted emotion values, so treat the field as a hint and audition it. Render the same line three times with a different guide each time, and keep the job ids so you can compare. Each request needs an Idempotency-Key, and the voice comes from an avatar handle you already own.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def wait(d):
    while not d["result_ready"]:
        if d.get("terminal"):
            raise RuntimeError(d.get("status"))
        time.sleep(d.get("next_poll_after_seconds") or 3)
        j = requests.get(d["status_url"], headers=H).json()
        d = j.get("data", j)
    j = requests.get(d["result_url"], headers=H).json()
    return j.get("data", j)["result"]

line = "We shipped the update this morning."
for guide in ["calm", "excited", "apologetic"]:
    body = {"transcript": line, "avatar_handle": "your_handle",
            "generation_config": {"emotion": guide, "speed": 1.0}}
    r = requests.post(API + "/v1/tts-1.0/generate", json=body,
                      headers={**H, "Idempotency-Key": "emotion-test-" + guide})
    r.raise_for_status()
    d = r.json()["data"]
    print(guide, d["job"]["id"])

How to judge the result

Listen blind. Have someone else shuffle the three files and name the emotion they hear. If listeners cannot tell calm from apologetic, the field is not doing what you need for that voice, and you should change the script wording instead of the guide.

A completed TTS job records model_id, voice, language, output_format, generation_config and speed as it was synthesized, per Jobs and results. Read those from the job before you render the next line so the take matches.

If you need the exact Microsoft voice, use Microsoft's own surfaces. If you need narration that lands inside a captioned, timed video, the Sume path is a TTS job, then Timeline 1.0 and video captions.

Sources

Related posts

More in Models

All Models posts

Written by Sume