Match voice emotion and music mood: one mood word, two Sume fields

Set generation_config.emotion on TTS and the emotion axis in the Music prompt from the same mood word, so a short's voice and bed do not argue.

5 min readSume
All posts

To keep a voiceover and a music bed in the same mood on Sume, write one mood phrase and paste it into both calls: generation_config.emotion on the text-to-speech job, and the emotion axis of the Music prompt. Sume does not link the two surfaces, so the match is yours to keep. A shared phrase is the cheapest way to stop a warm, slow narration from landing on a bright, driving bed.

Microsoft's MAI-Voice-2 page lists emotion tags as one of its features (read 2026-10-05). Tag-style control is a vendor design. On Sume, the TTS control is a plain string field, and the music control is a sentence in the prompt.

The two fields

The Music page asks for precise emotion words, and says darker choices are allowed: for example hushed and slightly melancholic, proud and nostalgic, or cocky and restless. It names seven axes in all. Emotion is the first one, followed by genre, tempo as a number, key and mode, instruments with texture, an arc with one named moment, and era or production.

Where the mood goes on each Sume surface (read 2026-10-05)
SurfaceFieldFormLimits in the docs or contract
TTS 1.0 (POST /v1/tts-1.0/generate)generation_config.emotionFree string, an emotion guideOptional. speed 0.6 to 1.5 and volume 0.5 to 2.0 sit beside it.
Music Router (POST /v1/music-router/generate)prompt, emotion axisA phrase such as hushed, slightly melancholic1 to 5000 characters. No seed, temperature or duration field.

A worked pairing

Take a 30-second product short with narration. Choose the mood phrase first: for example calm, warm, quietly confident. Put it in the TTS request as the emotion guide. Then open the music prompt with the same words: Calm, warm, quietly confident neo-soul bed, 78 BPM, F major. Rhodes through tape wow, soft sub bass, brushed snare. Instrumental, no vocals. No spoken word.

The Music page says to add no spoken word only under narration, and to end every prompt with the instrumental clause. Both clauses matter here: a bed that sings or speaks competes with the voice track you just generated.

Mixing the two

Timeline 1.0 takes the voice as the audio spine and the music as soundtrack. The soundtrack has gain_db, loop, fade_out_seconds up to 10, and duck_db from 0 to 20. duck_db needs a real spine, not audio.mode: silence. A mood match does not replace ducking: the bed still needs to sit under the voice.

When a scene should contrast with the one before it, the Music page suggests changing the broad genre family, moving the tempo at least 12 BPM, and changing the lead instrument. Then change the emotion phrase on the voice side in the same step, or the pair will drift.

A short checklist

  • Write the mood phrase once, in plain words, before any call.
  • Paste it into generation_config.emotion and into the first clause of the music prompt.
  • End the music prompt with the instrumental clause, and add the no-spoken-word clause under narration.
  • Listen to the generated audio. The Music page says the axes are creative directions, not guaranteed values.
  • Render with duck_db set, and check the voice stays on top.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume