Lower background music under narration: Timeline duck_db and gain_db

Keep a music bed from covering the voiceover in a Sume Timeline render: soundtrack gain_db, duck_db from 0 to 20, loop, and a fade-out up to 10 seconds.

4 min readSume
All posts

To lower background music under narration in a Sume render, add a soundtrack to the Timeline document with duck_db between 0 and 20, plus a gain_db for the bed's base level. Ducking needs a real audio spine, meaning an audio.url or audio.parts[] voiceover, so it is refused with duck_requires_audio_spine when audio.mode is silence. The docs list no separate price for ducking: the render stays $0.10 per output minute, rounded up.

The soundtrack fields and refusal codes come from Timeline 1.0, the price from API pricing, and the voice side from the Sume API reference, read on 2026-10-03. The docs do not say how loud the bed becomes when ducked beyond duck_db, so I give example values, not recommendations, and you should listen.

Which soundtrack fields matter?

All of them sit on one optional object, next to the voiceover spine.

Soundtrack fields in a Timeline 1.0 render, from the Timeline docs, read 2026-10-03.
FieldRange or ruleUse
urlSume-hosted audioThe music bed
gain_dbSet the base level of the bedTurn the bed down for the whole video
duck_db0 to 20Lower the bed while the spine speaks
looptrue or falseRepeat a short bed
fade_out_secondsUp to 10Ease the bed out at the end

How do I write the render body?

The body below puts a 30-second voiceover over one still, with a looped bed at -6 dB that ducks by 12 dB. The values are examples. Run POST /v1/timeline-1.0/plan first; it is unbilled and checks the document without creating a job, then send the render with an Idempotency-Key.

import json
body = {
    "audio": {"url": "https://media.sume.com/example/voiceover.wav", "duration_seconds": 30},
    "video": [{"source_url": "https://media.sume.com/example/still.png", "start": 0, "duration": 30}],
    "soundtrack": {
        "url": "https://media.sume.com/example/bed.mp3",
        "gain_db": -6,
        "duck_db": 12,
        "loop": True,
        "fade_out_seconds": 2,
    },
}
assert 0 <= body["soundtrack"]["duck_db"] <= 20
assert body["soundtrack"]["fade_out_seconds"] <= 10
print(json.dumps(body, indent=2))

What does the whole 30-second spot cost?

Add the three steps from the pricing page, read 2026-10-03. A 500-character script at $0.0475 per 1,000 characters is $0.02375. One music generation is $0.125. The Timeline render for a 30-second output rounds up to one minute, $0.10. The total is about $0.249, and ducking adds nothing the docs list. Swap in your own script length; the other two lines are flat.

What can go wrong?

Check the document before paying.

  • Fade longer than the video: a soundtrack fade past the output length fails as soundtrack_fade_exceeds_output.
  • A bed shorter than the voice needs loop: true, or it ends early.
  • If the voice itself is too quiet, fix it at the source with the TTS generation_config.volume, a multiplier from 0.5 to 2, before you duck the bed.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume