Duck a Music Bed Under a Voiceover in Timeline 1.0: $0.25 Total

Mix a generated music bed under a 45-second voiceover with Timeline 1.0: soundtrack duck_db, gain_db, loop and fade-out, plus the cost of music, TTS and render.

4 min readSume
All posts

The short answer

In Timeline 1.0, put the voiceover in audio as the spine and the music in soundtrack with duck_db between 0 and 20. The bed then drops by that many decibels while the spine speaks. A 45-second ad with a generated bed, a 600-character voiceover and one render costs about $0.254: music at $0.125 flat, TTS at $0.0285 and the render at $0.10 for one started minute.

The settings

The docs list these soundtrack fields: url, gain_db, loop, fade_out_seconds (up to 10) and duck_db (0 to 20). Ducking needs a real spine, not silence; duck_requires_audio_spine is the refusal. audio.gain_db runs from -60 to 12 and is not allowed with silence. Output fades fade_in_seconds and fade_out_seconds are 0 to 5 seconds each, and their sum cannot exceed the output length.

Cost of the 45-second ad (Sume catalog, read 2026-10-08)
ItemRateQuantityTotal
Music Router generation$0.125 flat1$0.125
Text to speech$0.0475 per 1,000 chars600 chars$0.0285
Timeline 1.0 render$0.10 per started minute45 s = 1$0.10
Total$0.2535

An example

The program below uses the voice as the spine for 45 seconds and ducks a looping bed by 12 dB with a 3-second fade at the end. The two video slots start at 0 and 20 seconds, and the second uses a half-second fade.

Music Router rejects a duration field, so ask for the length in the prompt, such as "a 30-second track, steady, no vocals". Because the bed has loop: true, a shorter track repeats to fill the ad. A hard loop point can be audible; ask the generator for a bed that fades or ends on a held note.

{
  "audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
             "duration_seconds": 45 },
  "soundtrack": { "url": "https://media.sume.com/artifacts/artf_demo/bed.mp3",
                  "gain_db": -6, "loop": true,
                  "duck_db": 12, "fade_out_seconds": 3 },
  "video": [
    { "source_url": "https://media.sume.com/artifacts/artf_demo/a.mp4", "start": 0, "duration": 20 },
    { "source_url": "https://media.sume.com/artifacts/artf_demo/b.mp4", "start": 20, "duration": 25,
      "transition": { "type": "fade", "duration": 0.5 } }
  ]
}

Mix checks

Play the ad on a phone speaker. A bed that sounds right on headphones can mask consonants on a small speaker, so raise duck_db in steps of 3. Keep the bed instrumental when the ad has speech. If you want to reuse the voiceover alone, join or split clips first with timeline audio, $0.01 per job.

Keep a log of the settings that worked, for example gain -6 dB and duck 12 dB, and reuse them across the campaign so ads sound alike. The whole mix is another render at $0.10 per started minute if you change it, so decide on gain and duck values on a 10-second test with audio.duration_seconds set to 10, which is also one billed minute. When the voice is a TTS job, record its voice id and settings from the job result so a fixed line can be remade in the same voice without remixing the whole bed. If the final cut is over 60 seconds, the render line grows by $0.10 per extra started minute, and the music bed must be long enough or must loop.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume