Duck music under dialogue and fade out a TikTok mini drama episode

Sume's Timeline render takes a soundtrack with duck_db (0-20) and a fade_out_seconds up to 5, so dialogue stays clear and each episode ends on a clean fade.

5 min readSume
All posts

To keep dialogue clear in a mini drama episode, send Timeline 1.0 a soundtrack with duck_db between 0 and 20, and set output.fade_out_seconds (0 to 5) to end on a fade. Ducking needs a real audio spine, so your dialogue track must be the spine and the music the bed. A 60-second render costs $0.10.

Both fields are in the Timeline 1.0 docs, and the render is plain ffmpeg on Sume's worker: no model call and no re-generation of any shot.

Build the spine first

Timeline wants one audio spine plus ordered video[] slots. The spine is your dialogue. If the voice lives inside a clip, pull it out with Audio detach ($0.01 per job): wav by default, mp3 at 128 kbps, channels: "mono" and 16 kHz if you want the STT shape. If you have several lines, audio.parts[] takes up to 20 gapless slices.

For a music bed, Music Router generates a track from a prompt (sume/music-auto, Lyria 3.5 today). It rejects duration, so state the length in the prompt, such as "a 60-second track, instrumental, no vocals".

The soundtrack block

soundtrack takes url, gain_db, loop, fade_out_seconds (up to 10) and duck_db (0 to 20). The docs say duck_db needs a real spine and not silence. Set loop: true if the bed is shorter than the episode, and output.fade_out_seconds for the whole mix and picture.

An audio.mode: "silence" spine is for declared length with no voice, and it cannot be ducked, so it suits a music-only teaser but not dialogue.

Timeline audio fields used for a mini drama (read 2026-10-07)
FieldRange / rule
audio.duration_seconds1-1800, sets output length
audio.parts[]Up to 20 gapless slices, no re-TTS
soundtrack.duck_db0-20, needs a real spine
soundtrack.fade_out_secondsUp to 10
output.fade_in_seconds / fade_out_seconds0-5 each, sum must fit the length
{
  "audio": {"url": "https://media.sume.com/artifacts/artf_demo/dialogue.wav", "duration_seconds": 58},
  "video": [
    {"source_url": "https://media.sume.com/artifacts/artf_demo/s1.mp4", "start": 0, "duration": 29},
    {"source_url": "https://media.sume.com/artifacts/artf_demo/s2.mp4", "start": 29, "duration": 29,
     "transition": {"type": "fade", "duration": 0.25}}
  ],
  "soundtrack": {"url": "https://media.sume.com/artifacts/artf_demo/bed.mp3", "gain_db": -6, "loop": true, "duck_db": 12},
  "output": {"fade_out_seconds": 1.5}
}

Where AI audio meets TikTok's label

TikTok's 2023 label announcement ties the creator label to realistic images, audio, or video that is completely generated or significantly edited by AI. An instrumental bed is not a realistic voice, but an AI voice in the dialogue is. If any of the mix is synthetic speech, switch on the AI-generated content toggle when you post.

Run POST /v1/timeline-1.0/plan first. It is unbilled and returns duration_seconds, segment_count, billable_minutes, and the estimate.

Test the mix on a phone

Most TikTok viewers listen on a phone speaker or earbuds, and a mix that sounds right on studio monitors can bury dialogue on a small speaker. Render a short test, play it on a phone at a low volume, and ask whether you can follow every line without leaning in.

If the dialogue is lost, increase duck_db a few decibels at a time within the 0-20 range, and remember that ducking needs a real audio spine to key from: a silent spine gives it nothing to react to. If the music feels like it disappears too hard, lower the value instead. The plan call is unbilled, so you can adjust settings freely, then pay only for the final render at $0.10 per output minute.

End the episode with a fade so the loop into the next video is not abrupt. output.fade_out_seconds accepts 0-5 seconds, and the soundtrack has its own fade_out_seconds up to 10. A short picture fade with a slightly longer music fade is a common, comfortable pairing.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume