The Patchbay_

Generate Music with MusicGen in Python

Meta's MusicGen is the most-used open text-to-music model. This is the full reference — both install routes, every generation parameter, melody conditioning, continuation, the 30-second limit and how to get past it, and the licensing catch.

Last updated 2026-09-12

MusicGen is Meta's text-to-music model, and the most widely used open instrumental generator. This page covers both ways to run it, every generation parameter that matters, how to get past the 30-second wall, melody conditioning, continuation, and the install problems that stop most people on their first attempt.

How MusicGen actually works

Understanding the architecture explains most of the parameters, so it's worth thirty seconds.

MusicGen doesn't generate a waveform directly. An audio codec called EnCodec compresses 32 kHz audio into a stream of discrete tokens at 50 tokens per second, across several parallel codebooks. A single autoregressive transformer then predicts those tokens, conditioned on a T5 embedding of your text prompt. EnCodec decodes the predicted tokens back into audio.

Three consequences follow directly, and they explain nearly every rough edge you'll hit:

Two ways to run it: audiocraft or transformers

This is the first real decision, and most tutorials never mention it.

 audiocraft (Meta)transformers (Hugging Face)
Installpip install audiocraftpip install transformers torch
Dependency pinsPins torch==2.1.0, documented for Python 3.9 — routinely fights modern environmentsTracks current torch and Python
Melody conditioningYes — generate_with_chromaVia the separate MusicgenMelody classes
>30 s generationBuilt in — automatic sliding windowNot built in; you roll it yourself
Loudness-normalised writingYes — audio_writeNo, save it yourself
MaintenanceEffectively frozenActively maintained

Pick audiocraft if you want melody conditioning or clips longer than 30 seconds without writing the stitching yourself. Pick transformers if you are putting MusicGen into an existing PyTorch project and don't want a pinned torch version dictating your whole dependency tree. The bulk of this guide uses audiocraft, with the transformers equivalent at the end.

Installing audiocraft without fighting it

AudioCraft's README asks for Python 3.9 and pins PyTorch 2.1.0. Installing it into an environment you care about is how people end up with a broken torch. Give it its own environment, always:

# a dedicated environment — do not install this next to anything you value
conda create -n musicgen python=3.9 -y
conda activate musicgen

# torch FIRST, matched to your CUDA version
python -m pip install 'torch==2.1.0'
python -m pip install setuptools wheel
python -m pip install -U audiocraft

Install torch before audiocraft. The package pulls in xformers, which compiles against whatever torch it finds, and letting pip resolve both at once is the single most common cause of install failure.

You also want ffmpeg on the system — audiocraft shells out to it for anything that isn't a WAV:

# macOS
brew install ffmpeg
# Debian/Ubuntu
sudo apt-get install ffmpeg
# or inside the conda env
conda install 'ffmpeg<5' -c conda-forge
Apple Silicon. MusicGen runs on Mac, but not comfortably. MPS support in the pinned torch version is incomplete for some ops, and you will likely need PYTORCH_ENABLE_MPS_FALLBACK=1, which silently moves the unsupported ops to CPU and erases most of the speed benefit. On an M-series machine, musicgen-small on CPU is the realistic path, at roughly real-time-times-ten for short clips. If Mac is your only hardware, ACE-Step is the better-behaved choice.

Choosing a checkpoint

Six base models, plus stereo fine-tunes of each. The VRAM column is inference at fp32 with a short duration; long generations and large batches push it higher.

CheckpointParamsVRAM (approx.)MelodyNotes
facebook/musicgen-small300M~4 GBNoThe only one that's tolerable on CPU
facebook/musicgen-medium1.5B~16 GBNoMeta's recommended quality/compute point
facebook/musicgen-large3.3B~24 GBNoBest text-only quality
facebook/musicgen-melody1.5B~16 GBYesMelody conditioning, the most interesting model here
facebook/musicgen-melody-large3.3B~24 GBYesMelody at large scale
facebook/musicgen-stereo-*as baseslightly higheras baseStereo fine-tunes of every model above

Meta's own guidance is that medium or melody is the sweet spot, and that a GPU with at least 16 GB is needed for the 1.5B models. Below that, you are on small, or generating very short clips.

The stereo models are genuine fine-tunes, not a post-processing widener: they emit two interleaved EnCodec streams. If the output is going anywhere near a mix, start with musicgen-stereo-medium rather than widening mono afterwards.

Your first generation

from audiocraft.models import MusicGen
from audiocraft.data.audio import audio_write

model = MusicGen.get_pretrained('facebook/musicgen-small')
model.set_generation_params(duration=8)

wav = model.generate(['upbeat 80s synthwave with a driving bassline'])

for i, one in enumerate(wav):
    audio_write(f'clip_{i}', one.cpu(), model.sample_rate,
                strategy='loudness', loudness_compressor=True)

print(model.sample_rate)   # 32000
print(model.audio_channels) # 1 for mono checkpoints

First run downloads several GB of weights to ~/.cache/huggingface. After that, load time is a few seconds.

audio_write adds the extension for you — pass 'clip_0', get clip_0.wav. Passing 'clip_0.wav' gets you clip_0.wav.wav, which catches everyone once.

Every generation parameter

These are the real defaults from MusicGen.set_generation_params. Most guides show you duration and stop.

ParameterDefaultWhat it does
use_samplingTrueSample from the distribution rather than taking the argmax. Setting this False gives greedy decoding, which sounds noticeably worse and repetitive. Leave it on.
top_k250Sample from the 250 most likely tokens. Lower (40–100) is more conservative and coherent; higher is more adventurous and more likely to wander.
top_p0.0Nucleus sampling. Zero means disabled, and top_k is used instead. Set it above 0 and it takes over. Try 0.9 if you want variety without the long tail.
temperature1.0Flattens or sharpens the distribution. Below 1.0 gives safer, more repetitive output; above 1.2 tends to disintegrate rhythmically.
duration30.0Seconds to generate. Above 30 triggers the sliding-window path below.
cfg_coef3.0Classifier-free guidance strength — how hard the model is pushed toward your prompt. This is the parameter most worth tuning. Raise to 4–5 to make it obey a complicated prompt; the cost is a thinner, more brittle sound. Below 2 it largely ignores you.
two_step_cfgFalseRuns conditional and unconditional passes separately instead of batched. Halves throughput, saves memory. A last resort when you're out of VRAM.
extend_stride18For generations past 30 s: how far the window advances each step. See below. Must be less than 30.

In practice, cfg_coef and top_k are the two you'll actually reach for. A reasonable starting point for something that needs to follow a detailed prompt:

model.set_generation_params(
    duration=20,
    cfg_coef=4.5,      # push harder toward the prompt
    top_k=100,         # tighter sampling, fewer odd detours
    temperature=0.95,  # marginally more conservative
)

Generating longer than 30 seconds

Ask for more than 30 seconds and audiocraft generates in overlapping windows: produce 30 s, advance by extend_stride, re-prime the model with the tail of what it just wrote, continue. The overlap is 30 − extend_stride seconds of context carried forward.

model.set_generation_params(
    duration=90,        # three windows
    extend_stride=15,   # 15 s of fresh audio per window, 15 s of overlap
)
wav = model.generate(['slow cinematic strings building to a climax'])

The trade-off is direct:

The honest limitation. This is a sliding window, not long-range composition. The model has no memory beyond its 30-second context, so a two-minute generation will not return to an earlier theme or resolve a structure — it will plausibly continue, and drift. For anything with real form, generate sections separately and arrange them in a DAW, or reach for a model built for full songs, like ACE-Step.

Melody conditioning

The melody models take a chroma representation of audio you supply and follow its harmonic contour while taking style from the text prompt. You hum the tune; it decides what the band sounds like.

import torchaudio
from audiocraft.models import MusicGen
from audiocraft.data.audio import audio_write

model = MusicGen.get_pretrained('facebook/musicgen-melody')
model.set_generation_params(duration=12, cfg_coef=3.0)

melody, sr = torchaudio.load('hummed_idea.wav')

wav = model.generate_with_chroma(
    descriptions=['lofi hip hop, warm rhodes, vinyl crackle, brushed drums'],
    melody_wavs=melody[None],   # note the batch dimension
    melody_sample_rate=sr,
    progress=True,
)
audio_write('with_melody', wav[0].cpu(), model.sample_rate, strategy='loudness')

Notes that matter in practice:

Continuing existing audio

Different from melody conditioning: generate_continuation takes real audio as a prompt and keeps playing from where it stopped, in the same material.

import torchaudio
from audiocraft.models import MusicGen
from audiocraft.data.audio import audio_write

model = MusicGen.get_pretrained('facebook/musicgen-medium')
model.set_generation_params(duration=20)

prompt, sr = torchaudio.load('four_bar_loop.wav')
prompt = prompt[..., : int(5 * sr)]   # first 5 seconds only

wav = model.generate_continuation(
    prompt=prompt[None],
    prompt_sample_rate=sr,
    descriptions=['keep the groove, add a saxophone solo'],
    progress=True,
)
audio_write('continued', wav[0].cpu(), model.sample_rate, strategy='loudness')

The prompt audio counts against the 30-second context. Feed it 20 seconds with duration=30 and you get 10 seconds of new music, not 30. Two to five seconds of prompt is usually enough to establish key and feel while leaving room to generate.

Batching, seeds and picking winners

MusicGen output is inconsistent enough that generating one clip and judging the model on it is a mistake. Generate several and audition. Multiple descriptions in one call are batched through the GPU together, which is far faster than looping:

prompts = [
    'dark techno, rolling 909 hats, sub bass, 130 bpm',
    'dark techno, rolling 909 hats, sub bass, 130 bpm',
    'dark techno, rolling 909 hats, sub bass, 130 bpm',
    'dark techno, distorted kick, acid 303 line, 132 bpm',
]
model.set_generation_params(duration=12)
wav = model.generate(prompts, progress=True)

for i, one in enumerate(wav):
    audio_write(f'take_{i:02d}', one.cpu(), model.sample_rate, strategy='loudness')

Pass the same string several times to get variations on one prompt. For reproducibility, seed torch before generating — audiocraft has no seed argument of its own:

import torch

torch.manual_seed(42)
wav = model.generate(['ambient drone, deep and slow'])

Writing files properly

audio_write does more than dump a tensor, and its defaults are not what you want for a set of clips you plan to compare:

ParameterDefaultWhat it does
format'wav'Also 'mp3', 'ogg', 'flac' — the non-WAV formats need ffmpeg
strategy'peak''peak', 'rms', 'loudness' or 'clip'. Use 'loudness' so clips are comparable by ear instead of by peak
loudness_headroom_db14Target headroom for the loudness strategy
loudness_compressorFalseSoft-limits before normalising; worth enabling for peaky material
mp3_rate320kbps, when format is mp3
add_suffixTrueAppends the extension — the source of the double-.wav problem
audio_write(
    'render',
    wav[0].cpu(),
    model.sample_rate,
    format='wav',
    strategy='loudness',
    loudness_headroom_db=14,
    loudness_compressor=True,
    add_suffix=True,      # writes render.wav
)

Judging generations by peak normalisation is misleading — a quiet, dense mix and a loud sparse one hit the same peak. Loudness normalisation makes an A/B actually meaningful.

Progress callbacks

Long generations look like a hang. Wire up a callback so they don't:

def on_progress(done, total):
    print(f'\r{done}/{total} tokens ({100 * done / total:.0f}%)', end='', flush=True)

model.set_custom_progress_callback(on_progress)
wav = model.generate(['epic orchestral trailer music'], progress=True)

The transformers route

If a pinned torch is unacceptable, the Hugging Face implementation avoids audiocraft entirely. Duration is expressed in tokens rather than seconds — at 50 tokens per second, max_new_tokens=256 is about 5.1 seconds:

import torch, scipy.io.wavfile
from transformers import AutoProcessor, MusicgenForConditionalGeneration

processor = AutoProcessor.from_pretrained('facebook/musicgen-small')
model = MusicgenForConditionalGeneration.from_pretrained(
    'facebook/musicgen-small', device_map='auto')

inputs = processor(
    text=['80s pop track with bassy drums and synth',
          '90s rock song with loud guitars and heavy drums'],
    padding=True,
    return_tensors='pt',
)

audio_values = model.generate(**inputs, do_sample=True,
                              guidance_scale=3, max_new_tokens=512)  # ~10 s

sr = model.config.audio_encoder.sampling_rate
scipy.io.wavfile.write('hf_out.wav', rate=sr, data=audio_values[0, 0].cpu().numpy())

guidance_scale here is the same concept as cfg_coef in audiocraft, with the same default of 3. Note there is no loudness normalisation and no built-in path past 30 seconds — you get a raw tensor and the rest is yours.

Writing prompts that work

MusicGen was trained on text-audio pairs from stock music libraries, which shapes what it responds to. Prompts that read like a stock-library description outperform poetic ones.

More on this in the prompt guide.

Troubleshooting

SymptomCause and fix
CUDA out of memoryDrop to a smaller checkpoint, shorten duration, reduce batch size, or set two_step_cfg=True. Memory scales with all three of model size, duration and batch.
Cannot stride by more than max generation durationextend_stride is ≥ 30. Set it lower.
xformers build fails on installtorch wasn't installed first, or your Python is too new for the pinned version. Fresh environment, torch first.
Output is silence or noiseUsually a mangled melody tensor — check the shape is (batch, channels, samples) and that melody[None] is there.
NotImplementedError on MPSAn op has no Metal kernel. Set PYTORCH_ENABLE_MPS_FALLBACK=1, or just use CPU.
ffmpeg not found when writing mp3Install ffmpeg system-wide, or write WAV and convert afterwards.
Every clip sounds identicalA seed is pinned somewhere, or use_sampling got set to False.

Licensing: the part that matters commercially

The code and the weights have different licences. AudioCraft's code is MIT. The MusicGen weights are CC-BY-NC 4.0 — non-commercial. That covers the model, and you should assume it colours what you do with the output. If you are building anything commercial, this is a blocker, and no amount of post-processing changes it. Verify the current terms on the model card before relying on any of this.

The training data was licensed — Meta's own collection, Shutterstock and Pond5, about 20,000 hours — so the provenance is cleaner than some models. That doesn't loosen the weights licence.

If you need permissive terms, the practical open alternatives are ACE-Step (Apache-2.0) and Stable Audio Open (Stability AI Community License, free below a revenue threshold). Neither matches MusicGen on instrumental quality, and both are usable.

Where MusicGen actually fits

After all that, an honest summary. MusicGen is very good at eight to thirty seconds of instrumental texture in a recognisable genre, and at following a melody you give it. It is not good at song structure, vocals, long-form coherence, or clean high-end. Treat it as a sketch generator and a sample source — generate a lot, keep the 5% that's usable, and build around it in a DAW.

For that last part, pedalboard handles post-processing in the same Python script, and librosa will tell you what tempo and key you actually got, which is rarely what you asked for.

Run Stable Audio Open locallyStereo 44.1 kHz audio and SFX, with a commercially usable licence.Generate full songs with ACE-StepApache-2.0, does vocals, and far faster than MusicGen.MusicGen alternativesHow it compares to everything else in the directory.

← AI music generation hub