MAI-Voice-2.1 emotion control vs Sume's emotion field
Microsoft lists emotion control on MAI-Voice-2.1. Sume's TTS takes a free-text emotion string, speed and volume. What each gives you, and how to test it.

Does Sume's text to speech have the emotion control Microsoft announced for MAI-Voice-2.1? It has an emotion control, but not the same one and not that model. Sume's TTS request takes generation_config.emotion, a free-text guide of up to 64 characters, plus a speed multiplier and a volume multiplier. Microsoft's model is not in Sume's TTS Router, whose catalog is Sonic only.
Microsoft's MAI-Voice-2.1 page (read 2026-10-04) says both MAI-Voice-2.1 and MAI-Voice-2.1-Flash include granular emotion control and zero-shot voice prompting. It does not describe the control surface in the part of the page we could read, so this post compares only what each side states.
What each side documents
The Sume column below comes from the TTS 1.0 request contract in the API reference and its OpenAPI snapshot.
| Control | MAI-Voice-2.1 (Microsoft page) | Sume TTS 1.0 / Router |
|---|---|---|
| Emotion | Granular emotion control, no parameter list on the page | generation_config.emotion, string, 1-64 characters, described as an emotion guide |
| Speed | Not stated on the page | generation_config.speed, 0.6 to 1.5 |
| Volume | Not stated on the page | generation_config.volume, 0.5 to 2.0 |
| Voice from a short clip | Instant voice matching from a reference clip | Voice chosen by voice.id or an avatar reference; no clone endpoint in this contract |
Try an emotion string
The contract does not enumerate accepted emotion values, so treat the field as a hint and audition it. Render the same line three times with a different guide each time, and keep the job ids so you can compare. Each request needs an Idempotency-Key, and the voice comes from an avatar handle you already own.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def wait(d):
while not d["result_ready"]:
if d.get("terminal"):
raise RuntimeError(d.get("status"))
time.sleep(d.get("next_poll_after_seconds") or 3)
j = requests.get(d["status_url"], headers=H).json()
d = j.get("data", j)
j = requests.get(d["result_url"], headers=H).json()
return j.get("data", j)["result"]
line = "We shipped the update this morning."
for guide in ["calm", "excited", "apologetic"]:
body = {"transcript": line, "avatar_handle": "your_handle",
"generation_config": {"emotion": guide, "speed": 1.0}}
r = requests.post(API + "/v1/tts-1.0/generate", json=body,
headers={**H, "Idempotency-Key": "emotion-test-" + guide})
r.raise_for_status()
d = r.json()["data"]
print(guide, d["job"]["id"])
How to judge the result
Listen blind. Have someone else shuffle the three files and name the emotion they hear. If listeners cannot tell calm from apologetic, the field is not doing what you need for that voice, and you should change the script wording instead of the guide.
A completed TTS job records model_id, voice, language, output_format, generation_config and speed as it was synthesized, per Jobs and results. Read those from the job before you render the next line so the take matches.
If you need the exact Microsoft voice, use Microsoft's own surfaces. If you need narration that lands inside a captioned, timed video, the Sume path is a TTS job, then Timeline 1.0 and video captions.
Sources
Related posts
More in Models
- MiniMax H3 Max: the prompt-adherence variant on Sume
fal describes MiniMax H3 Max as tuned for prompt adherence. On Sume, minimax-h3-max runs 480p to 1080p for 5 to 15 s with frames and references.
- MiniMax H3 limits: 9 images, 3 videos, 3 audio, file caps
MiniMax's H3 guide caps prompts at 7,000 characters and references at 9 images, 3 videos and 3 audio files. Cheat sheet with the Sume limits beside it.
- Nano Banana reference image limits: Lite, 2 and Pro compared
Nano Banana 2 Lite takes up to 14 object images, Nano Banana 2 takes 10 object, 4 character and 3 style, Pro takes 6 object and 5 character.
- Pixal3D multi-view: prepare the input images
Pixal3D added multi-view inference in September 2026 under an MIT license. Sume has no 3D, but reference-image edit can prepare input views.
Written by Sume