Lower background music under narration: Timeline duck_db and gain_db
Keep a music bed from covering the voiceover in a Sume Timeline render: soundtrack gain_db, duck_db from 0 to 20, loop, and a fade-out up to 10 seconds.

To lower background music under narration in a Sume render, add a soundtrack to the Timeline document with duck_db between 0 and 20, plus a gain_db for the bed's base level. Ducking needs a real audio spine, meaning an audio.url or audio.parts[] voiceover, so it is refused with duck_requires_audio_spine when audio.mode is silence. The docs list no separate price for ducking: the render stays $0.10 per output minute, rounded up.
The soundtrack fields and refusal codes come from Timeline 1.0, the price from API pricing, and the voice side from the Sume API reference, read on 2026-10-03. The docs do not say how loud the bed becomes when ducked beyond duck_db, so I give example values, not recommendations, and you should listen.
Which soundtrack fields matter?
All of them sit on one optional object, next to the voiceover spine.
| Field | Range or rule | Use |
|---|---|---|
url | Sume-hosted audio | The music bed |
gain_db | Set the base level of the bed | Turn the bed down for the whole video |
duck_db | 0 to 20 | Lower the bed while the spine speaks |
loop | true or false | Repeat a short bed |
fade_out_seconds | Up to 10 | Ease the bed out at the end |
How do I write the render body?
The body below puts a 30-second voiceover over one still, with a looped bed at -6 dB that ducks by 12 dB. The values are examples. Run POST /v1/timeline-1.0/plan first; it is unbilled and checks the document without creating a job, then send the render with an Idempotency-Key.
import json
body = {
"audio": {"url": "https://media.sume.com/example/voiceover.wav", "duration_seconds": 30},
"video": [{"source_url": "https://media.sume.com/example/still.png", "start": 0, "duration": 30}],
"soundtrack": {
"url": "https://media.sume.com/example/bed.mp3",
"gain_db": -6,
"duck_db": 12,
"loop": True,
"fade_out_seconds": 2,
},
}
assert 0 <= body["soundtrack"]["duck_db"] <= 20
assert body["soundtrack"]["fade_out_seconds"] <= 10
print(json.dumps(body, indent=2))What does the whole 30-second spot cost?
Add the three steps from the pricing page, read 2026-10-03. A 500-character script at $0.0475 per 1,000 characters is $0.02375. One music generation is $0.125. The Timeline render for a 30-second output rounds up to one minute, $0.10. The total is about $0.249, and ducking adds nothing the docs list. Swap in your own script length; the other two lines are flat.
What can go wrong?
Check the document before paying.
- Fade longer than the video: a soundtrack fade past the output length fails as
soundtrack_fade_exceeds_output. - A bed shorter than the voice needs
loop: true, or it ends early. - If the voice itself is too quiet, fix it at the source with the TTS
generation_config.volume, a multiplier from 0.5 to 2, before you duck the bed.
Sources
Related posts
More in Developers
- LRC lyrics file from an AI music track: Sume STT segments in Python
YouTube accepts .lrc lyric files. A short Python script turns STT sentence segments from a generated song into [mm:ss.xx] lines. Check results before upload.
- Luma Agents API polling: 20 s wait, 10-minute video timeout, and Sume
Luma's quickstart says wait 20 seconds, then poll, with a hard 2-minute image and 10-minute video timeout, not backoff. A matching Sume loop that actually runs.
- Luma Ray 2 to Ray 3.2 migration: the field map and what changes
Luma is retiring Ray 3, Ray 2, Ray 2 Flash and more for ray-3.2 on one /v1/generations endpoint. The field map, the poll-only change, and Sume's contrast.
- macos-14 brownouts from Oct 5: rerun a Sume submit with the same key
GitHub's macos-14 runners fail on purpose in brownout windows before the 2026-11-02 retirement. Derive a stable Idempotency-Key so a re-run does not bill twice.
Written by Sume