Music bed under narration: the no-spoken-word clause for Lyria 3.5

To get a music bed that stays out of a voiceover's way, end the Sume music prompt with no vocals, no spoken word, then duck it under the voice in Timeline.

4 min readSume
All posts

To get a music bed that stays out of a voiceover's way, close the music prompt with Instrumental, no vocals, no spoken word, and keep the arrangement sparse in the range where the voice sits. Sume's Music docs say to add no spoken word only when a narration will sit on top, because a track with its own vocal line fights the voice.

Lyria 3.5 is the engine behind the Sume Music Router today, and Google's Gemini API changelog (read 2026-10-02) says it generates full-length songs with duration and structure control. That is more than a bed needs, so you ask for less.

What do I write in the prompt?

Name the job in the brief: an underscore for narration. Give a tempo, a key, two to four instruments, and an arc that stays low. Then close with the clause. Do not put exclusions in a negative_prompt field; a non-empty value returns HTTP 400 with public_reason negative_prompt_unsupported.

Prompt choices for a narrated bed. Sume Music 1.0 docs, read 2026-10-02.
ChoiceDoAvoid
Closing clauseInstrumental, no vocals, no spoken wordLeaving vocals open
ExclusionsWrite them in the positive promptA non-empty negative_prompt
LengthDescribe the length and arc in the brief (a direction, not a guarantee)duration or duration_seconds fields
ArcOne named moment, low in the mixDense full-band peaks under speech

How do I sit it under the voice?

Use Timeline 1.0. It accepts an optional soundtrack with url, gain_db, loop, fade_out_seconds up to 10 and duck_db from 0 to 20. Ducking needs a real audio spine, meaning your narration, not silence. Set duck_db so the music dips while the voice speaks, and set a fade out so the bed does not stop abruptly.

Do I need to generate the exact length?

No. The soundtrack field has loop and a fade, and the Timeline render length is set by the spine. Generate a track a bit longer than the narration if you can, and let the fade handle the end.

What does the bed cost?

Music 1.0 lists $0.125 per accepted generation; Timeline render pricing is on the Timeline page.

How do I check the bed against the voice?

Render a short test with the real narration, not a placeholder. Listen on laptop speakers and on a phone, since a bed that sounds fine in headphones can mask consonants on a small speaker. If the voice is hard to follow, lower the soundtrack gain_db before you regenerate anything; a gain change only needs a new Timeline render, while a new generation is another $0.125 charge.

If the problem is the arrangement itself, for example a lead instrument that sits in the same range as the voice, change the instruments in the brief. Name two or three quiet textures, such as soft pads or a felt piano, and leave out a melodic lead. Because there is no seed, expect a different take each time, and keep the file you accept.

When should the bed have no melody at all?

For explainer and tutorial narration, a pulse and a pad are usually enough. Write that in the brief: a steady pulse at a stated tempo, a warm pad, no lead melody, instrumental, no vocals, no spoken word. For a product reveal with a pause in the speech, allow one named moment, such as a rise at a stated timestamp, and say where the narration is silent.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume