Separate Stems with Demucs
Split any track into vocals, drums, bass and other — then do something useful with the result. Model selection, quality flags, memory tuning, the Python API, and the artefacts you should expect.
Last updated 2026-09-12
Demucs splits a finished stereo mix back into its parts — vocals, drums, bass and everything else. It is MIT licensed, it runs on a laptop, and it is comfortably the best open source separator available.
Most guides stop at the one-line command. This one covers model selection, the quality flags that actually change the result, memory tuning, the Python API, and what to do about the artefacts you will definitely hear.
What separation is for
- Remixing and sampling — isolate a drum break or a vocal to build something new around it.
- Backing tracks and practice — pull the guitar out to play along, or keep only the drums.
- Feeding other models. This is the underrated one. A clean isolated stem is dramatically better input for pitch detection, transcription or melody conditioning than a full mix. MusicGen's melody mode works far better on a separated stem, and Basic Pitch transcribes an isolated instrument far more accurately than a mix.
- Repair. Separate, fix one element, recombine.
- Karaoke and dialogue isolation — the obvious ones.
Install
pip install -U demucs
# or, with uv, no install at all
uvx demucs my_track.mp3
# ffmpeg for non-WAV input and output
brew install ffmpeg # macOS
sudo apt-get install ffmpeg # Debian/Ubuntu
Python 3.10 or later. ffmpeg is optional but you want it — without it you are limited to WAV in and out.
The one-liner
demucs my_track.mp3
Output lands in separated/htdemucs/<track name>/ as four WAVs: vocals.wav, drums.wav, bass.wav, other.wav.
That's the whole basic workflow. Everything below is about making it better.
Choosing a model
SDR — signal-to-distortion ratio, in dB — is the standard separation quality metric. The difference between 7.7 and 9.0 dB is clearly audible.
| Model | SDR | Stems | Notes |
|---|---|---|---|
| htdemucs | 9.0 dB | 4 | The default. Hybrid Transformer Demucs. Start here |
| htdemucs_ft | 9.0 dB | 4 | Fine-tuned per source. Around 4× slower for a modest improvement — worth it for final work, not for triage |
| htdemucs_6s | — | 6 | Adds guitar and piano. The piano stem is acknowledged as weak; guitar is genuinely useful |
| hdemucs_mmi | 7.7 dB | 4 | Hybrid Demucs v3. Faster, lighter, noticeably worse |
| mdx | — | 4 | MDX challenge winner, trained on MusDB HQ only |
| mdx_extra | — | 4 | Trained with extra data including the test set — strong in practice, not comparable on benchmarks |
| mdx_q / mdx_extra_q | — | 4 | Quantised: much smaller download, slightly worse output. For constrained machines |
demucs -n htdemucs_ft my_track.wav # best 4-stem quality
demucs -n htdemucs_6s my_track.wav # adds guitar and piano
demucs -n mdx_extra_q my_track.wav # small and fast
Practical advice: use htdemucs to audition, then re-run the keepers through htdemucs_ft. Paying 4× the compute on tracks you end up discarding is the most common waste of time here.
The flags that change quality
| Flag | Default | What it does |
|---|---|---|
| -n MODEL | htdemucs | Model selection |
| --shifts N | 1 | Shift-trick averaging: run the model N times on randomly time-shifted copies and average. Worth up to about 0.2 dB SDR, costs N× the time. Use 2–5 for final renders |
| --overlap F | 0.25 | Overlap between processing windows. Raising to 0.5 reduces seam artefacts at moderate cost |
| --segment S | model default | Window length in seconds. Lower to fit less memory. Transformer models cap around 7.8 s |
| --two-stems STEM | off | Produce just that stem and everything else. Roughly halves write time and disk |
| -d DEVICE | cuda if available | cuda, cpu, or mps on Apple Silicon |
| -j N | 0 | Parallel jobs. Real speedup on multi-core CPU, at proportional memory cost |
| --mp3 | off | Write MP3 instead of WAV; --mp3-bitrate defaults to 320 |
| --flac | off | FLAC output; needs ffmpeg |
| --int24 / --float32 | 16-bit | Output bit depth. Use --float32 if stems feed further processing |
| -o DIR | separated | Output directory |
| --filename | template | Naming pattern for outputs |
A quality-first invocation for something you actually care about:
demucs -n htdemucs_ft \
--shifts 5 \
--overlap 0.5 \
--float32 \
-o stems \
my_track.wav
That will be slow — several times real time even on a GPU. It is also as good as open source separation currently gets.
Karaoke and stem-only extraction
# vocals.wav + no_vocals.wav
demucs --two-stems=vocals my_track.mp3
# isolate the drums, everything else in one file
demucs --two-stems=drums my_track.mp3
# karaoke track straight to mp3
demucs --two-stems=vocals --mp3 --mp3-bitrate 320 my_track.mp3
--two-stems accepts any stem the model produces — vocals, drums, bass, other, plus guitar and piano on the 6-source model.
Memory and speed
Demucs needs about 3 GB of GPU memory minimum and is comfortable around 7 GB. On CPU, expect processing time of roughly 1.5× the track length per core-ish, which makes a full album an overnight job.
# smaller windows = less memory
demucs --segment 10 my_track.wav
# CPU with 8 parallel jobs
demucs -d cpu -j 8 my_track.wav
# Apple Silicon
demucs -d mps my_track.wav
# last resort on a tight GPU
PYTORCH_NO_CUDA_MEMORY_CACHING=1 demucs --segment 10 my_track.wav
If you are hitting OOM, work down this list in order:
- Lower --segment. This is the direct memory dial. Do not go below 10 s unless you have to; very short segments hurt quality.
- Drop --shifts back to 1 and --overlap to 0.25.
- Set PYTORCH_NO_CUDA_MEMORY_CACHING=1, which trades speed for a smaller footprint.
- Switch to a quantised model (mdx_q).
- Fall back to -d cpu and accept the wait.
The Python API
Two options. The quick one just calls the CLI:
import demucs.separate
demucs.separate.main([
'--mp3', '--two-stems', 'vocals', '-n', 'mdx_extra', 'my_track.mp3'
])
That is fine for scripting, but it shells out and writes files. For real integration use the Separator class, which returns tensors:
from demucs.api import Separator, save_audio
separator = Separator(
model='htdemucs_ft',
device='cuda', # 'cpu' or 'mps'
shifts=2,
overlap=0.35,
segment=None, # model default
jobs=0,
progress=True,
)
origin, stems = separator.separate_audio_file('my_track.wav')
for name, source in stems.items():
save_audio(source, f'{name}.wav', samplerate=separator.samplerate)
drums = stems['drums'] # torch tensor, (channels, samples)
print(drums.shape, separator.samplerate)
separate_audio_file returns a tuple of the original waveform and a dict keyed by stem name. Nothing touches disk unless you call save_audio, which makes it straightforward to pipe stems into analysis or generation without a round trip through the filesystem.
Progress callbacks
For anything with a UI, the callback gives you progress out of a long-running separation:
from demucs.api import Separator
def on_progress(data):
total = data['audio_length']
done = data['segment_offset']
print(f"\r{data['state']} {100 * done / total:.1f}%", end='', flush=True)
separator = Separator(model='htdemucs', callback=on_progress)
origin, stems = separator.separate_audio_file('my_track.wav')
Raising KeyboardInterrupt inside the callback aborts the separation cleanly — that's the documented cancellation mechanism.
Batch processing a library
import pathlib
from demucs.api import Separator, save_audio
separator = Separator(model='htdemucs', device='cuda', progress=False)
src = pathlib.Path('library')
out = pathlib.Path('stems')
for track in sorted(src.glob('**/*.flac')):
dest = out / track.stem
if (dest / 'vocals.wav').exists():
continue # resumable
dest.mkdir(parents=True, exist_ok=True)
print('separating', track.name)
_, stems = separator.separate_audio_file(track)
for name, source in stems.items():
save_audio(source, dest / f'{name}.wav', samplerate=separator.samplerate)
Stems into other tools
Separation is most valuable as the first step of a pipeline. Two that work well:
Stem to MIDI. Transcription on a full mix is poor; on an isolated stem it's usable:
# 1. isolate the instrument
demucs --two-stems=other track.wav -o work
# 2. transcribe the clean stem
basic-pitch ./midi work/htdemucs/track/other.wav
Stem as a melody reference. MusicGen's melody conditioning uses chroma, which a dense mix confuses. Feed it a separated stem instead:
import torchaudio
from demucs.api import Separator
from audiocraft.models import MusicGen
from audiocraft.data.audio import audio_write
separator = Separator(model='htdemucs')
_, stems = separator.separate_audio_file('reference.wav')
melody = stems['other'] # the harmonic stem
model = MusicGen.get_pretrained('facebook/musicgen-melody')
model.set_generation_params(duration=15)
wav = model.generate_with_chroma(
['orchestral arrangement, warm strings, slow build'],
melody[None],
separator.samplerate,
)
audio_write('rearranged', wav[0].cpu(), model.sample_rate, strategy='loudness')
Living with the artefacts
You will hear problems. Knowing which are fixable saves time:
- Metallic ringing on vocals — phase reconstruction error. More --shifts and higher --overlap reduce it. It never fully disappears.
- Cymbals bleeding into "other" — cymbal transients are broadband and genuinely ambiguous. Expected behaviour.
- Bass leaking into drums — kick and bass overlap in frequency and time. Worst on heavily compressed masters.
- Reverb tails in the wrong stem — the model assigns the tail to whichever source it best matches, which often isn't where you want it. Not fixable through flags.
- Ghost vocals in "other" — usually backing vocals or vocal-like synths. htdemucs_ft handles these better.
- Worse results on loud masters — heavy limiting destroys the dynamic cues the model relies on. If you have access to a pre-master, separate that instead.
General rule: sparse, dynamic, well-recorded material separates well. Dense, loud, heavily processed material does not, and no flag combination rescues it.
Demucs or Spleeter?
Spleeter came first and is still widely referenced. Demucs is better in essentially every way that matters — substantially higher SDR, better artefact behaviour, actively maintained. Spleeter's remaining advantage is speed on CPU and a smaller install.
Use Demucs unless you are processing at volume on CPU-only hardware where speed beats quality. The full comparison is at Spleeter vs Demucs.
Licensing
Demucs is MIT — code and models — which is about as unencumbered as it gets. There is no non-commercial clause and no revenue threshold.
Where it fits
Demucs is the most immediately useful model in this entire section. It solves a concrete problem, it runs on ordinary hardware, the licence is clean, and the output is good enough to use in real work.
Treat it as infrastructure: the step that turns an existing recording into material other tools can work on. Pair it with librosa for analysis, Basic Pitch for transcription, or pedalboard for processing — all in the same Python script.
How this model’s licence compares with every other open model → · GPU & VRAM requirements →