Match voice emotion and music mood: one mood word, two Sume fields
Set generation_config.emotion on TTS and the emotion axis in the Music prompt from the same mood word, so a short's voice and bed do not argue.

To keep a voiceover and a music bed in the same mood on Sume, write one mood phrase and paste it into both calls: generation_config.emotion on the text-to-speech job, and the emotion axis of the Music prompt. Sume does not link the two surfaces, so the match is yours to keep. A shared phrase is the cheapest way to stop a warm, slow narration from landing on a bright, driving bed.
Microsoft's MAI-Voice-2 page lists emotion tags as one of its features (read 2026-10-05). Tag-style control is a vendor design. On Sume, the TTS control is a plain string field, and the music control is a sentence in the prompt.
The two fields
The Music page asks for precise emotion words, and says darker choices are allowed: for example hushed and slightly melancholic, proud and nostalgic, or cocky and restless. It names seven axes in all. Emotion is the first one, followed by genre, tempo as a number, key and mode, instruments with texture, an arc with one named moment, and era or production.
| Surface | Field | Form | Limits in the docs or contract |
|---|---|---|---|
| TTS 1.0 (POST /v1/tts-1.0/generate) | generation_config.emotion | Free string, an emotion guide | Optional. speed 0.6 to 1.5 and volume 0.5 to 2.0 sit beside it. |
| Music Router (POST /v1/music-router/generate) | prompt, emotion axis | A phrase such as hushed, slightly melancholic | 1 to 5000 characters. No seed, temperature or duration field. |
A worked pairing
Take a 30-second product short with narration. Choose the mood phrase first: for example calm, warm, quietly confident. Put it in the TTS request as the emotion guide. Then open the music prompt with the same words: Calm, warm, quietly confident neo-soul bed, 78 BPM, F major. Rhodes through tape wow, soft sub bass, brushed snare. Instrumental, no vocals. No spoken word.
The Music page says to add no spoken word only under narration, and to end every prompt with the instrumental clause. Both clauses matter here: a bed that sings or speaks competes with the voice track you just generated.
Mixing the two
Timeline 1.0 takes the voice as the audio spine and the music as soundtrack. The soundtrack has gain_db, loop, fade_out_seconds up to 10, and duck_db from 0 to 20. duck_db needs a real spine, not audio.mode: silence. A mood match does not replace ducking: the bed still needs to sit under the voice.
When a scene should contrast with the one before it, the Music page suggests changing the broad genre family, moving the tempo at least 12 BPM, and changing the lead instrument. Then change the emotion phrase on the voice side in the same step, or the pair will drift.
A short checklist
- Write the mood phrase once, in plain words, before any call.
- Paste it into
generation_config.emotionand into the first clause of the music prompt. - End the music prompt with the instrumental clause, and add the no-spoken-word clause under narration.
- Listen to the generated audio. The Music page says the axes are creative directions, not guaranteed values.
- Render with
duck_dbset, and check the voice stays on top.
Sources
Related posts
More in Media tools
- Microsoft's Content Provenance Detection: what you can check on a file
Foundry has a detection website and API for provenance. What the page says it checks, its limits, and how to use it on an AI clip or image from any generator.
- Microsoft: C2PA may not survive a crop, transcode or compression
Microsoft's provenance page lists the edits that can drop credentials and watermarks. Which of them a Sume trim, filter or caption job could be, and a check.
- Fit TTS audio to the 5 to 14.8 second H3 Max lip-sync window
Sume's H3 Max lip-sync accepts audio of 5 to 14.8 seconds and never clamps. Plan scripts of about 63 to 197 characters, and send anything else to Fabric.
- MiniMax H3 takes 9 image references: a slot plan for one product clip
Hailuo 3.0 (MiniMax H3) accepts up to 9 reference images per clip. A 9-slot plan for a product ad, and the request to send it on Sume.
Written by Sume