Music bed for a SaaS product demo: a Lyria brief and ducking
A product demo needs a quiet instrumental bed under narration. A seven-axis Lyria 3.5 brief, a Timeline soundtrack with duck_db, and what it costs.

For a product-demo video, ask the Music Router for a steady, sparse instrumental in a fixed tempo, then place it as a Timeline soundtrack with duck_db so it dips under the narration. Write the brief on seven axes, end it with "Instrumental, no vocals.", and state the length in words. One generation costs a fixed $0.125, and a one-minute render adds $0.10.
The brief
The Music 1.0 page recommends a scene-specific brief with seven axes. For a demo, choose calm over clever: a clear tempo, few instruments, and an arc that stays out of the narration's way. Example prompt: "Focused, quietly optimistic tech-pop, 96 BPM, A minor. Soft Rhodes pulse, muted sub bass, light shaker, one glassy pad. Sparse for the first minute with a slight lift at 0:40, nothing in the vocal range. 2020s production, dry and close. A 90-second track. Instrumental, no vocals."
- Emotion: named precisely (focused, quietly optimistic).
- Tempo as a number and key: so it is repeatable.
- Two to four instruments with texture.
- One named moment, not a dramatic build.
- Length written into the prompt, because
durationis rejected.
The render
Put the narration as the audio spine and the music as soundtrack. Timeline 1.0 accepts soundtrack with url, gain_db, loop, fade_out_seconds up to 10 and duck_db from 0 to 20. duck_db needs a real spine, so it works under a voiceover and is refused when the spine is silence (duck_requires_audio_spine). Start with duck_db around 10 and adjust by ear.
| Step | Surface | Price |
|---|---|---|
| Generate music | Music Router | $0.125 per generation |
| Render 60 s with soundtrack | Timeline 1.0 | $0.10 per output minute, rounded up |
| Total, one take | Music plus render | $0.225 |
Check before you publish
Call the unbilled POST /v1/timeline-1.0/plan first; it returns billable_minutes and estimated_cost_usd_micros without creating a job. Then listen to the render at the quietest speaker you expect viewers to use, and confirm the narration is clear. The plan call cannot predict warnings about padded or looped sources, which appear on the finished job.
If the track is shorter than the video, set loop or ask for a longer one in the prompt. For results and polling, see Jobs and results.
Sources
Related posts
More in Use cases
- IPTC compositeSynthetic vs compositeWithTrainedAlgorithmicMedia
Both IPTC terms mean a composite with generative AI in it. One says edited using generative AI; the other says a composite where at least one element is Gen AI.
- IPTC minorHumanEdits was retired: use humanEdits instead
IPTC retired minorHumanEdits on 2024-09-17 and replaced it with humanEdits. Which Sume edits (trim, filter, caption) are non-generative ffmpeg passes.
- Keep one character consistent in a 30-second Seedance 2.5 clip
Character consistency in a single 30 s Seedance 2.5 pass: send reference_image_urls on Sume, test at 480p, and judge stills. No guarantee, a cheap test plan.
- Kling 4.0 keyframes at 0, 6, 12, 20, 28 s: a 30-second ad on Sume
Kling's guide suggests keyframes near 0, 6, 12, 20 and 28 seconds for a product ad. How to build the same plan on Sume with first-and-last-frame clips.
Written by Sume