The Patchbay_

Separate Stems with Demucs

Split any track into vocals, drums, bass and other — then do something useful with the result. Model selection, quality flags, memory tuning, the Python API, and the artefacts you should expect.

Last updated 2026-09-12

Demucs splits a finished stereo mix back into its parts — vocals, drums, bass and everything else. It is MIT licensed, it runs on a laptop, and it is comfortably the best open source separator available.

Most guides stop at the one-line command. This one covers model selection, the quality flags that actually change the result, memory tuning, the Python API, and what to do about the artefacts you will definitely hear.

What separation is for

What it is not. Separation is reconstruction, not un-mixing. There is no hidden multitrack inside a stereo file; the model is estimating what each source probably sounded like. Expect artefacts, especially on dense or heavily limited masters, and expect that the sum of the stems is close to but not identical to the original.

Install

pip install -U demucs

# or, with uv, no install at all
uvx demucs my_track.mp3

# ffmpeg for non-WAV input and output
brew install ffmpeg          # macOS
sudo apt-get install ffmpeg  # Debian/Ubuntu

Python 3.10 or later. ffmpeg is optional but you want it — without it you are limited to WAV in and out.

The one-liner

demucs my_track.mp3

Output lands in separated/htdemucs/<track name>/ as four WAVs: vocals.wav, drums.wav, bass.wav, other.wav.

That's the whole basic workflow. Everything below is about making it better.

Choosing a model

SDR — signal-to-distortion ratio, in dB — is the standard separation quality metric. The difference between 7.7 and 9.0 dB is clearly audible.

ModelSDRStemsNotes
htdemucs9.0 dB4The default. Hybrid Transformer Demucs. Start here
htdemucs_ft9.0 dB4Fine-tuned per source. Around 4× slower for a modest improvement — worth it for final work, not for triage
htdemucs_6s—6Adds guitar and piano. The piano stem is acknowledged as weak; guitar is genuinely useful
hdemucs_mmi7.7 dB4Hybrid Demucs v3. Faster, lighter, noticeably worse
mdx—4MDX challenge winner, trained on MusDB HQ only
mdx_extra—4Trained with extra data including the test set — strong in practice, not comparable on benchmarks
mdx_q / mdx_extra_q—4Quantised: much smaller download, slightly worse output. For constrained machines
demucs -n htdemucs_ft my_track.wav     # best 4-stem quality
demucs -n htdemucs_6s my_track.wav     # adds guitar and piano
demucs -n mdx_extra_q my_track.wav     # small and fast

Practical advice: use htdemucs to audition, then re-run the keepers through htdemucs_ft. Paying 4× the compute on tracks you end up discarding is the most common waste of time here.

The flags that change quality

FlagDefaultWhat it does
-n MODELhtdemucsModel selection
--shifts N1Shift-trick averaging: run the model N times on randomly time-shifted copies and average. Worth up to about 0.2 dB SDR, costs N× the time. Use 2–5 for final renders
--overlap F0.25Overlap between processing windows. Raising to 0.5 reduces seam artefacts at moderate cost
--segment Smodel defaultWindow length in seconds. Lower to fit less memory. Transformer models cap around 7.8 s
--two-stems STEMoffProduce just that stem and everything else. Roughly halves write time and disk
-d DEVICEcuda if availablecuda, cpu, or mps on Apple Silicon
-j N0Parallel jobs. Real speedup on multi-core CPU, at proportional memory cost
--mp3offWrite MP3 instead of WAV; --mp3-bitrate defaults to 320
--flacoffFLAC output; needs ffmpeg
--int24 / --float3216-bitOutput bit depth. Use --float32 if stems feed further processing
-o DIRseparatedOutput directory
--filenametemplateNaming pattern for outputs

A quality-first invocation for something you actually care about:

demucs -n htdemucs_ft \
       --shifts 5 \
       --overlap 0.5 \
       --float32 \
       -o stems \
       my_track.wav

That will be slow — several times real time even on a GPU. It is also as good as open source separation currently gets.

Karaoke and stem-only extraction

# vocals.wav + no_vocals.wav
demucs --two-stems=vocals my_track.mp3

# isolate the drums, everything else in one file
demucs --two-stems=drums my_track.mp3

# karaoke track straight to mp3
demucs --two-stems=vocals --mp3 --mp3-bitrate 320 my_track.mp3

--two-stems accepts any stem the model produces — vocals, drums, bass, other, plus guitar and piano on the 6-source model.

Memory and speed

Demucs needs about 3 GB of GPU memory minimum and is comfortable around 7 GB. On CPU, expect processing time of roughly 1.5× the track length per core-ish, which makes a full album an overnight job.

# smaller windows = less memory
demucs --segment 10 my_track.wav

# CPU with 8 parallel jobs
demucs -d cpu -j 8 my_track.wav

# Apple Silicon
demucs -d mps my_track.wav

# last resort on a tight GPU
PYTORCH_NO_CUDA_MEMORY_CACHING=1 demucs --segment 10 my_track.wav

If you are hitting OOM, work down this list in order:

  1. Lower --segment. This is the direct memory dial. Do not go below 10 s unless you have to; very short segments hurt quality.
  2. Drop --shifts back to 1 and --overlap to 0.25.
  3. Set PYTORCH_NO_CUDA_MEMORY_CACHING=1, which trades speed for a smaller footprint.
  4. Switch to a quantised model (mdx_q).
  5. Fall back to -d cpu and accept the wait.
Apple Silicon. -d mps works and is substantially faster than CPU. It is still slower than a mid-range NVIDIA GPU, and occasionally an op falls back to CPU. For a handful of tracks it's perfectly usable.

The Python API

Two options. The quick one just calls the CLI:

import demucs.separate

demucs.separate.main([
    '--mp3', '--two-stems', 'vocals', '-n', 'mdx_extra', 'my_track.mp3'
])

That is fine for scripting, but it shells out and writes files. For real integration use the Separator class, which returns tensors:

from demucs.api import Separator, save_audio

separator = Separator(
    model='htdemucs_ft',
    device='cuda',       # 'cpu' or 'mps'
    shifts=2,
    overlap=0.35,
    segment=None,        # model default
    jobs=0,
    progress=True,
)

origin, stems = separator.separate_audio_file('my_track.wav')

for name, source in stems.items():
    save_audio(source, f'{name}.wav', samplerate=separator.samplerate)

drums = stems['drums']          # torch tensor, (channels, samples)
print(drums.shape, separator.samplerate)

separate_audio_file returns a tuple of the original waveform and a dict keyed by stem name. Nothing touches disk unless you call save_audio, which makes it straightforward to pipe stems into analysis or generation without a round trip through the filesystem.

Progress callbacks

For anything with a UI, the callback gives you progress out of a long-running separation:

from demucs.api import Separator

def on_progress(data):
    total = data['audio_length']
    done = data['segment_offset']
    print(f"\r{data['state']}  {100 * done / total:.1f}%", end='', flush=True)

separator = Separator(model='htdemucs', callback=on_progress)
origin, stems = separator.separate_audio_file('my_track.wav')

Raising KeyboardInterrupt inside the callback aborts the separation cleanly — that's the documented cancellation mechanism.

Batch processing a library

import pathlib
from demucs.api import Separator, save_audio

separator = Separator(model='htdemucs', device='cuda', progress=False)

src = pathlib.Path('library')
out = pathlib.Path('stems')

for track in sorted(src.glob('**/*.flac')):
    dest = out / track.stem
    if (dest / 'vocals.wav').exists():
        continue                      # resumable
    dest.mkdir(parents=True, exist_ok=True)
    print('separating', track.name)
    _, stems = separator.separate_audio_file(track)
    for name, source in stems.items():
        save_audio(source, dest / f'{name}.wav', samplerate=separator.samplerate)

Stems into other tools

Separation is most valuable as the first step of a pipeline. Two that work well:

Stem to MIDI. Transcription on a full mix is poor; on an isolated stem it's usable:

# 1. isolate the instrument
demucs --two-stems=other track.wav -o work

# 2. transcribe the clean stem
basic-pitch ./midi work/htdemucs/track/other.wav

Stem as a melody reference. MusicGen's melody conditioning uses chroma, which a dense mix confuses. Feed it a separated stem instead:

import torchaudio
from demucs.api import Separator
from audiocraft.models import MusicGen
from audiocraft.data.audio import audio_write

separator = Separator(model='htdemucs')
_, stems = separator.separate_audio_file('reference.wav')
melody = stems['other']                       # the harmonic stem

model = MusicGen.get_pretrained('facebook/musicgen-melody')
model.set_generation_params(duration=15)

wav = model.generate_with_chroma(
    ['orchestral arrangement, warm strings, slow build'],
    melody[None],
    separator.samplerate,
)
audio_write('rearranged', wav[0].cpu(), model.sample_rate, strategy='loudness')

Living with the artefacts

You will hear problems. Knowing which are fixable saves time:

General rule: sparse, dynamic, well-recorded material separates well. Dense, loud, heavily processed material does not, and no flag combination rescues it.

Demucs or Spleeter?

Spleeter came first and is still widely referenced. Demucs is better in essentially every way that matters — substantially higher SDR, better artefact behaviour, actively maintained. Spleeter's remaining advantage is speed on CPU and a smaller install.

Use Demucs unless you are processing at volume on CPU-only hardware where speed beats quality. The full comparison is at Spleeter vs Demucs.

Licensing

Demucs is MIT — code and models — which is about as unencumbered as it gets. There is no non-commercial clause and no revenue threshold.

The licence covers the tool, not the music. Separating a copyrighted recording produces a derivative work of that recording. Demucs's MIT licence has nothing to say about it. Practising along to a stem is one thing; releasing one is another, and the rights holder's position is unchanged by how you extracted it.

Where it fits

Demucs is the most immediately useful model in this entire section. It solves a concrete problem, it runs on ordinary hardware, the licence is clean, and the output is good enough to use in real work.

Treat it as infrastructure: the step that turns an existing recording into material other tools can work on. Pair it with librosa for analysis, Basic Pitch for transcription, or pedalboard for processing — all in the same Python script.

Generate music with MusicGenUse a separated stem as a melody reference.Spleeter vs DemucsThe two open separators, compared in detail.Basic PitchTurn an isolated stem into MIDI.

← AI music generation hub