Demucs is archived: still use it to split a song into stems?
Demucs was archived on January 1, 2025 but is MIT-licensed and still splits drums, bass, vocals and other. How to feed it a track from Sume audio detach.

Yes, Demucs still separates a song into stems, but nobody maintains it. The repository says "this repository is not maintained anymore" and was archived on January 1, 2025, with the author's fork saying the same. The code is MIT-licensed, installs with python3 -m pip install -U demucs, and the default model returns four stems: drums, bass, vocals and other.
What does Demucs output?
The README lists several models. The default is htdemucs, the Hybrid Transformer version; htdemucs_ft is fine-tuned, slower and potentially better; a 6-source experimental model adds guitar and piano, though the README says the piano source is not working great.
| Model | Sources | Note |
|---|---|---|
| htdemucs | drums, bass, vocals, other | Default, newest Hybrid Transformer |
| htdemucs_ft | same four | Fine-tuned, slower, potentially better |
| 6-source (experimental) | adds guitar and piano | Piano not working great |
| mdx, mdx_extra, hdemucs_mmi | varies | Older or competition models |
What hardware does it need?
For GPU use the README asks for a minimum of 3 GB of VRAM and recommends 7 GB. On CPU, processing takes roughly 1.5 times the track length. A --segment parameter helps when memory is tight.
How do you get a track from a video to Demucs?
If the song lives inside a Sume-hosted video, audio detach turns that video's audio into a wav (the default, pcm_s16le) or a 128 kbps mp3, at $0.01 per job. The source can be up to 1800 seconds and the output up to 900, so a longer file needs a range. Download the wav and run demucs on it locally. Detach does not separate vocals from music; it extracts the whole track.
A rule of thumb: keep wav for anything you will process again, since the docs describe it as sample-exact.
Does Sume return stems when it makes music?
Not in the docs we read. A Music 1.0 job returns one audio artifact, usually audio/mpeg, and the Music Router reads the same way: one artifact per generation, priced at the fixed Music price, with no stem field. If you need stems, ask for a bed without vocals in the prompt ("Instrumental, no vocals.") instead of splitting afterwards, or run Demucs over the finished track yourself.
An archived tool is a risk you can accept for a one-off cleanup, and a poor fit for a production pipeline you must support for years. Pin the version and keep the model files.
Where does stem separation help a video project?
The common cases are ducking by hand and cleanup. Dropping the vocals stem lets you put a voiceover over a song without fighting the singer. Dropping the drums can calm a beat that fights a title card. For a licensed track you own, separation is just another editing step. For a track you do not have rights to, separation does not change that.
- Voiceover over a song you have rights to: keep drums, bass and other; drop vocals.
- A karaoke style cut: keep vocals only.
- An archive clip with music under speech: results vary, so audition the stems before trusting them.
What are the risks of an archived tool?
Nobody will fix a bug or update a dependency for you. Python and PyTorch move on, and a pin that works today can fail on a new machine next year. Mitigate it by recording the exact version, saving the model weights you downloaded and keeping a container or lockfile. For a one-off stem job none of this matters; for a service that must run for years it is the main cost.
The README also points to a fork by the original author that carries the same not-maintained notice, so switching to it does not change the support picture.
Sources
Related posts
More in Media tools
- Descript AI music and sound effects vs generating music by API
Descript's 2026-09-17 update generates music and sound effects on request. Sume's Music Router generates music from a prompt; it has no sound-effect route.
- Descript's ElevenLabs voice library vs Sume TTS voices
Descript's voice picker now uses the ElevenLabs voice library. Sume text-to-speech uses its own voices, and blocks a voice-language mismatch with a 409.
- Descript scrubbing 2-5x faster vs a video-frames contact sheet
Descript's 2026-09-17 update speeds scrubbing with pre-computed video chunks. For a quick look at a file, Sume video-frames returns 1 to 24 stills.
- Dim bright footage before burning Reel captions: video filter 0.45
A dim op multiplies luma by an amount between 0 and 1 for $0.02, and the check call is free. Use it to keep captions legible on bright clips, then caption.
Written by Sume