Stable Audio web app mixing vs Sume timeline soundtrack gain and duck
Stable Audio's web app has level, pan, mute and solo per track. Sume's timeline soundtrack has gain, ducking, loop and fade-out, but no pan, mute or solo.

Stability's Stable Audio web app lets you shape a generated piece with per-track level, pan, mute and solo, effects and track extension, then bounce it. Sume's mixing is narrower: a timeline soundtrack has gain_db, duck_db, loop and fade_out_seconds, and the docs list no pan, mute or solo. Use the web app for hands-on mixing; use Sume when the mix has to be a repeatable request.
Stability's details come from Sharing a new way to work with Stable Audio, read on 2026-10-02. Sume's come from the Timeline docs.
What can you do in the Stable Audio web app?
The post describes prompt-based direction, an audio-to-audio mode, level, pan, mute and solo controls, per-track and master effects, track extension and a bounce step. It is in beta at StableAudio.com. The post does not say how many tracks a project holds or what the bounce formats are, so those are left out.
What mixing controls does a Sume timeline have?
In a timeline request, audio.gain_db accepts -60 to 12 dB. The soundtrack object takes a url, gain_db, loop, fade_out_seconds up to 10 and duck_db from 0 to 20, which lowers the bed while speech plays. Separate Timeline audio jobs concat or split up to 20 parts, as WAV by default or MP3, for $0.01 a job on files up to 1,800 seconds.
A request that puts a track under a video with ducking looks like this in outline; check the Timeline docs for the full clip schema before sending.
# soundtrack fragment of a timeline request body
{
"soundtrack": {
"url": "https://media.sume.com/example/bed.mp3",
"gain_db": -12,
"loop": true,
"duck_db": 10,
"fade_out_seconds": 3
}
}How do the mixing controls line up?
| Control | Stable Audio web app | Sume timeline |
|---|---|---|
| Level | Per-track level | audio.gain_db -60..12; soundtrack.gain_db |
| Pan | Yes | Not documented |
| Mute and solo | Yes | Not documented |
| Ducking under speech | Not mentioned in the post | duck_db 0..20 |
| Loop and fade-out | Not mentioned in the post | loop, fade_out_seconds up to 10 |
| Effects | Per-track and master | Not documented |
When is ducking better than a manual mute?
A mute or solo is a decision you make by ear on a finished arrangement. Ducking is a rule: while the voice track plays, the bed drops by duck_db decibels, and it returns when speech stops. For a narrated video made by a script, a rule is the only thing that works, because no person is there to automate the fader.
The Timeline docs cap duck_db at 20 and fade_out_seconds at 10. If you need a deeper cut or a longer tail, bake it into the bed file beforehand and keep the Sume values modest.
Gain is the other half. soundtrack.gain_db sets the bed's base level, and audio.gain_db (-60 to 12) sets the level of the clip audio itself, so the voice-to-music balance is two numbers you can store with the request and reproduce exactly.
What can you not do on either side?
Be explicit about the gaps before you commit to a workflow. A mix that needs panning or solo has to be finished elsewhere, and a render that must be reproduced from a request has to be expressed in Sume's fields.
- Sume does not render a stereo pan, solo a stem or apply effects chains.
- The web app's post does not describe a request API, so a script cannot drive it from that page.
- Neither page promises loudness normalization; measure the final file.
What should you do about missing controls?
Do the creative mix where the controls exist, then bring the bounced file into Sume as the soundtrack url. Sume will set level, loop it under the video and duck it for speech. Do not expect Sume to pan a stem or solo a part, and do not expect the web app to render a repeatable video request. Related: mix voice with background music.
Sources
Related posts
More in Comparisons
- StepAudio 3 Gen: voice, SFX and music in one clip vs Sume jobs
StepFun's stepaudio-3-gen-preview makes voice, effects, ambience and music in one audio output. Sume uses separate music, speech and Timeline mix steps.
- StepAudio 3 Music from a dry vocal or reference audio vs Sume
StepAudio 3 Music accepts lyrics, vocals or reference audio, even scoring a dry vocal. Sume Music takes a text prompt and one optional image, no audio.
- Synthesia burned-in captions on dubbed videos, and Sume captions
Synthesia added one-click burned-in captions to dubbed videos on 9/30/2026. Here is what that means, and how to burn captions onto a finished video URL on Sume.
- Which Synthesia plan includes API access? Pro is limited
Synthesia lists no API on Basic or Starter, a limited API with 360 minutes a year on Pro, and full access on Enterprise. What a Sume API key gives you instead.
Written by Sume