What is audio ducking? Meaning, settings, and how it works
Audio ducking automatically lowers music while a voice speaks and raises it in the pauses. How the sidechain works, and what each setting does.

Audio ducking is an automatic volume dip: one track, such as background music, gets quieter while another, such as a voice, is sounding, then comes back up in the gaps. The music ducks out of the way of the speech, which is where the name comes from.
It is done with a sidechain compressor: a compressor on the music whose level detector listens to the voice instead. The settings below are from the FFmpeg filter documentation for its sidechaincompress filter, and the Sume values from Timeline 1.0, the Sume API reference, and Sume's code, read on 2026-09-28.
How does audio ducking work?
A normal compressor turns a track down when that track gets loud. A sidechain compressor turns one track down when a second track gets loud. FFmpeg's documentation describes its filter this way: it takes two inputs and returns one, and the first input is processed depending on the second input's signal. For ducking, the first input is the music and the second, the key, is the voice.
When the voice rises above a threshold, the compressor starts pulling the music down; when the voice falls below it again, the music is released back to its level. A few settings decide how that feels.
What do the ducking settings mean?
These are the controls a sidechain compressor has. The FFmpeg column is that filter's default; the Sume column is what a Timeline 1.0 render uses in current code.
| Setting | What it controls | FFmpeg default | Sume render |
|---|---|---|---|
| Threshold | The key level above which the music starts to duck | 0.125 | 0.01 (−40 dBFS) |
| Ratio | How hard the music is pulled down past the threshold | 2 (range 1–20) | Derived from duck_db |
| Attack | Milliseconds the key has to rise above the threshold before the dip starts | 20 ms | 20 ms |
| Release | Milliseconds the key has to fall below the threshold before the dip eases | 250 ms | 300 ms |
| Detection | Peak or RMS level of the key | RMS | RMS |
Is ducking the same as turning the music down?
No. Turning the music down sets one lower level for the whole track, so it stays quiet even when nobody is speaking. Ducking moves: the music sits at its normal level, drops under each phrase, and rises in the pauses. You can use both: a bed level plus a duck.
Sume's Timeline 1.0 render has both knobs on its soundtrack: gain_db sets the bed level, and duck_db, 0–20, is the target dip while the voice spine is speaking; omit duck_db and the bed stays static. Add background music to a video with an API shows the request and how its duck behaves, and How to mix voice with background music covers levels.
How do I duck music under a voice with FFmpeg?
FFmpeg's documentation gives this example: the first input is compressed depending on the signal of the second, and the compressed signal is then merged with the second input. Put the music first and the voice second, and add an output file name at the end:
One catch in FFmpeg's example: its documentation describes amerge as merging audio streams into a single multi-channel stream. The amix filter instead mixes multiple audio inputs into a single output, and amix is what Sume's render uses after the duck in current code.
ffmpeg -i main.flac -i sidechain.flac -filter_complex "[1:a]asplit=2[sc][mix];[0:a][sc]sidechaincompress[compr];[compr][mix]amerge"Sources
Related posts
More in Media tools
- What size is 16:9 in pixels? Common 16:9 and 9:16 sizes
16:9 is a shape, not one size: 1280 × 720, 1920 × 1080, 2560 × 1440 and 3840 × 2160 are all 16:9. How to work out any size, and how to make one.
- Twitter video ad specs: X Ads sizes, length, and file rules
X recommends video ads of 15 seconds or less, up to 1 GB (ideally under 30 MB), H.264 at 29.97 or 30 fps, in six sizes from 9:16 to 1.91:1.
- YouTube Masthead specs: video size, autoplay, and captions
A YouTube Masthead plays a Public or Unlisted YouTube video, ideally 16:9 at 1920×1080 or higher, and autoplays muted for up to 30 s on desktop.
- YouTube multi-language audio: add dubbed tracks, made with AI
YouTube multi-language audio lets one video carry dubbed tracks you upload yourself. The rules, the upload steps, and how to make each track with AI.
Written by Sume