Music bed under narration: the no-spoken-word clause for Lyria 3.5
To get a music bed that stays out of a voiceover's way, end the Sume music prompt with no vocals, no spoken word, then duck it under the voice in Timeline.

To get a music bed that stays out of a voiceover's way, close the music prompt with Instrumental, no vocals, no spoken word, and keep the arrangement sparse in the range where the voice sits. Sume's Music docs say to add no spoken word only when a narration will sit on top, because a track with its own vocal line fights the voice.
Lyria 3.5 is the engine behind the Sume Music Router today, and Google's Gemini API changelog (read 2026-10-02) says it generates full-length songs with duration and structure control. That is more than a bed needs, so you ask for less.
What do I write in the prompt?
Name the job in the brief: an underscore for narration. Give a tempo, a key, two to four instruments, and an arc that stays low. Then close with the clause. Do not put exclusions in a negative_prompt field; a non-empty value returns HTTP 400 with public_reason negative_prompt_unsupported.
| Choice | Do | Avoid |
|---|---|---|
| Closing clause | Instrumental, no vocals, no spoken word | Leaving vocals open |
| Exclusions | Write them in the positive prompt | A non-empty negative_prompt |
| Length | Describe the length and arc in the brief (a direction, not a guarantee) | duration or duration_seconds fields |
| Arc | One named moment, low in the mix | Dense full-band peaks under speech |
How do I sit it under the voice?
Use Timeline 1.0. It accepts an optional soundtrack with url, gain_db, loop, fade_out_seconds up to 10 and duck_db from 0 to 20. Ducking needs a real audio spine, meaning your narration, not silence. Set duck_db so the music dips while the voice speaks, and set a fade out so the bed does not stop abruptly.
Do I need to generate the exact length?
No. The soundtrack field has loop and a fade, and the Timeline render length is set by the spine. Generate a track a bit longer than the narration if you can, and let the fade handle the end.
What does the bed cost?
Music 1.0 lists $0.125 per accepted generation; Timeline render pricing is on the Timeline page.
How do I check the bed against the voice?
Render a short test with the real narration, not a placeholder. Listen on laptop speakers and on a phone, since a bed that sounds fine in headphones can mask consonants on a small speaker. If the voice is hard to follow, lower the soundtrack gain_db before you regenerate anything; a gain change only needs a new Timeline render, while a new generation is another $0.125 charge.
If the problem is the arrangement itself, for example a lead instrument that sits in the same range as the voice, change the instruments in the brief. Name two or three quiet textures, such as soft pads or a felt piano, and leave out a melodic lead. Because there is no seed, expect a different take each time, and keep the file you accept.
When should the bed have no melody at all?
For explainer and tutorial narration, a pulse and a pad are usually enough. Write that in the brief: a steady pulse at a stated tempo, a warm pad, no lead melody, instrumental, no vocals, no spoken word. For a product reveal with a pause in the speech, allow one named moment, such as a rise at a stated timestamp, and say where the narration is silent.
Sources
Related posts
More in Use cases
- Score a scene from its still: image_url on Sume music, Lyria 3.5
Pass the accepted scene still as image_url on a Sume music request so the score matches the picture. What the field accepts, what it does not, and the cost.
- Translate text inside an image: one Nano Banana 2 edit per language
Google says Nano Banana 2 can translate and localize text within an image. Loop one edit per language through Sume POST /v1/images and keep the layout fixed.
- Nano Banana Pro 4:5 is not 1080x1350: crop and resize in Python
On Sume, Nano Banana Pro takes 4:5 at about 928x1152, not exact 1080x1350. Resize to Instagram's size locally with Pillow, with the small crop explained.
- Narrate a 2,000-word blog post with an AI voice: cost and steps
A 2,000-word post is roughly 12,000 characters, so one Sume TTS job at $0.0475 per 1,000 characters. The steps, settings and what to check before publishing.
Written by Sume