What is audio ducking? Meaning, settings, and how it works

Audio ducking automatically lowers music while a voice speaks and raises it in the pauses. How the sidechain works, and what each setting does.

4 min readSume
All posts

Audio ducking is an automatic volume dip: one track, such as background music, gets quieter while another, such as a voice, is sounding, then comes back up in the gaps. The music ducks out of the way of the speech, which is where the name comes from.

It is done with a sidechain compressor: a compressor on the music whose level detector listens to the voice instead. The settings below are from the FFmpeg filter documentation for its sidechaincompress filter, and the Sume values from Timeline 1.0, the Sume API reference, and Sume's code, read on 2026-09-28.

How does audio ducking work?

A normal compressor turns a track down when that track gets loud. A sidechain compressor turns one track down when a second track gets loud. FFmpeg's documentation describes its filter this way: it takes two inputs and returns one, and the first input is processed depending on the second input's signal. For ducking, the first input is the music and the second, the key, is the voice.

When the voice rises above a threshold, the compressor starts pulling the music down; when the voice falls below it again, the music is released back to its level. A few settings decide how that feels.

What do the ducking settings mean?

These are the controls a sidechain compressor has. The FFmpeg column is that filter's default; the Sume column is what a Timeline 1.0 render uses in current code.

From the FFmpeg filter documentation (read 2026-09-28) and Sume's Timeline 1.0 docs and code.
SettingWhat it controlsFFmpeg defaultSume render
ThresholdThe key level above which the music starts to duck0.1250.01 (−40 dBFS)
RatioHow hard the music is pulled down past the threshold2 (range 1–20)Derived from duck_db
AttackMilliseconds the key has to rise above the threshold before the dip starts20 ms20 ms
ReleaseMilliseconds the key has to fall below the threshold before the dip eases250 ms300 ms
DetectionPeak or RMS level of the keyRMSRMS

Is ducking the same as turning the music down?

No. Turning the music down sets one lower level for the whole track, so it stays quiet even when nobody is speaking. Ducking moves: the music sits at its normal level, drops under each phrase, and rises in the pauses. You can use both: a bed level plus a duck.

Sume's Timeline 1.0 render has both knobs on its soundtrack: gain_db sets the bed level, and duck_db, 0–20, is the target dip while the voice spine is speaking; omit duck_db and the bed stays static. Add background music to a video with an API shows the request and how its duck behaves, and How to mix voice with background music covers levels.

How do I duck music under a voice with FFmpeg?

FFmpeg's documentation gives this example: the first input is compressed depending on the signal of the second, and the compressed signal is then merged with the second input. Put the music first and the voice second, and add an output file name at the end:

One catch in FFmpeg's example: its documentation describes amerge as merging audio streams into a single multi-channel stream. The amix filter instead mixes multiple audio inputs into a single output, and amix is what Sume's render uses after the duck in current code.

ffmpeg -i main.flac -i sidechain.flac -filter_complex "[1:a]asplit=2[sc][mix];[0:a][sc]sidechaincompress[compr];[compr][mix]amerge"

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume