Duck music under a voiceover on Sume Timeline: duck_db, gain, fade

Sume Timeline 1.0 soundtracks take gain_db, loop, fade_out_seconds up to 10 and duck_db from 0 to 20. Settings for a voiceover over a music bed.

4 min readSume
All posts

On a Sume Timeline 1.0 render, the voiceover is the audio spine and the music is the soundtrack. Set the soundtrack's duck_db (0 to 20) so the music dips while the voice speaks, gain_db to set its level, fade_out_seconds (up to 10) for the ending, and loop if the track is shorter than the video.

Timeline renders cost $0.10 per output minute, rounded up, with a spine of 1 to 1,800 seconds.

The soundtrack controls

From the Timeline 1.0 docs and reference.

Timeline 1.0 soundtrack and spine facts, from Sume docs (read 2026-10-07)
ControlRange or ruleUse
duck_db0 to 20How far the music drops under the spine
gain_dbA level in dBOverall music level
loopOn or offRepeat a short track under a longer video
fade_out_secondsUp to 10Fade the track at the end
Spine length1 to 1,800 secondsThe voiceover or master audio
Spine under 32 kHzFlagged audio_spine_low_fidelityUse 44.1 or 48 kHz voice for video
Video slots1 to 200, at least one requiredThe picture

A starting recipe

These are choices, not guarantees. Listen on speakers and headphones.

  • Generate the voice as wav or mp3 at 44100 Hz or above with Sume TTS.
  • Generate a music bed from a prompt that asks for an instrumental with space for a voiceover.
  • Start duck_db at a moderate value and raise it until the words are clear without the music vanishing.
  • Set fade_out_seconds to match the last shot, up to 10 seconds.
  • Turn on loop only if the music is shorter than the picture.

Check the mix before you publish

Play the render once on a phone speaker, because that is where most short videos are heard. If the words are hard to follow, raise duck_db or ask for a sparser music prompt. If the music disappears completely, lower duck_db. Re-rendering costs another $0.10 per output minute, so settle the voice and music files first, then tune the mix in as few renders as you can.

Cost of a 30-second spot

A 600-character voiceover is 3 cents on TTS. A music bed is $0.125 per generation. A 30-second render is one output minute at $0.10. Those three lines total $0.255, before any video clips you generate for the picture, which are priced by their own models. Add video captions at $0.20 per job if you want them.

If your voice is delivered as several takes, join them first with timeline audio at $0.01, or pass them as audio parts to the render and skip the extra job.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume