Music bed shorter than the voiceover: loop, fade and duck in Timeline
If a Music 1.0 bed ends before your narration, Sume's Timeline soundtrack can loop it, fade out up to 10 seconds and duck 0 to 20 dB under speech. The settings.

When a Music 1.0 bed ends before the voiceover, set loop on the Timeline soundtrack. Timeline will repeat the bed to the length of the audio spine, apply a fade-out of up to 10 seconds, and duck the bed by 0 to 20 dB under the narration. The spine is your voiceover; the soundtrack is the music.
The soundtrack fields
Timeline 1.0 takes one audio spine (your narration, up to 1,800 seconds) plus an optional soundtrack bed. The bed sits under the spine. These are the soundtrack options the docs list.
| Field | Range or type | What it does |
|---|---|---|
| url | Sume-hosted audio | The bed file |
| gain_db | number | Bed level |
| loop | boolean | Repeat the bed to fill the spine |
| fade_out_seconds | 0 to 10 | Fade at the end of the render |
| duck_db | 0 to 20 | Lower the bed under the spine; needs a real spine, not silence |
Loop or regenerate
Music 1.0 is a flat $0.125 per generation regardless of length, so there is no price reason to loop a short bed. The reason is control: there is no duration parameter, you steer length with the prompt ('a 2-minute track'), and the result can come back shorter or longer than you asked. A 3-minute narration with a 2-minute bed is a loop, not a failure.
If you want a bed that never repeats, write section markers such as [0:00-0:30] Intro in the prompt and ask for the length you need. If a repeat is fine, loop and save a retry.
Errors to expect
Timeline rejects a fade longer than the output with soundtrack_fade_exceeds_output, and duck_db on a silent spine with duck_requires_audio_spine. The plan call (unbilled) runs the same schema and compiler checks, so run it before you render. The render is $0.10 per ceil(output minute): a 3-minute narration is $0.30, and the bed adds the single $0.125 generation.
A starting point for speech
A bed under narration usually needs a gentle level. Start with duck_db around 10 and a fade_out_seconds of 4, then listen. These are starting values to tune by ear, not Sume recommendations. Keep the prompt for the bed instrumental, with no vocals and no spoken word, so it does not fight the voice.
Cost of the whole stack
For a 3-minute narrated clip: narration of about 2,700 characters at the planning assumption of 900 characters a minute is 2,700 x 47.5 = 128,250 micro-dollars, 13 cents; one bed is $0.125; the render is 3 x $0.10 = $0.30. Total about $0.555. Looping the bed adds nothing, since loop is a render setting.
The planning figure of 900 characters a minute is an assumption for this estimate, not a Sume number.
Sources
Related posts
More in Media tools
- Music API duration_seconds rejected: set track length in the prompt
Sume Music 1.0 and the Music Router reject duration and duration_seconds. Put the length in the prompt (a 2-minute track) and use time markers.
- Music from a still: image_url conditioning costs the same $0.125
Sume Music accepts an optional public HTTPS image_url as conditioning. The price stays a fixed $0.125 whether or not you send the image.
- Music from a still: image_url conditioning, still $0.125
The Music Router accepts an optional public HTTPS image_url to condition the track on a still. Request example, keeping a score consistent, fixed price.
- Music prompt limit: 5000 characters is plenty for a Lyria brief
The Sume Music prompt takes 1 to 5000 characters. A full seven-axis brief with section markers is far shorter, so the limit rarely binds.
Written by Sume