Generate Music with MusicGen in Python
Meta's MusicGen is the most-used open text-to-music model. This is the full reference — both install routes, every generation parameter, melody conditioning, continuation, the 30-second limit and how to get past it, and the licensing catch.
Last updated 2026-09-12
MusicGen is Meta's text-to-music model, and the most widely used open instrumental generator. This page covers both ways to run it, every generation parameter that matters, how to get past the 30-second wall, melody conditioning, continuation, and the install problems that stop most people on their first attempt.
How MusicGen actually works
Understanding the architecture explains most of the parameters, so it's worth thirty seconds.
MusicGen doesn't generate a waveform directly. An audio codec called EnCodec compresses 32 kHz audio into a stream of discrete tokens at 50 tokens per second, across several parallel codebooks. A single autoregressive transformer then predicts those tokens, conditioned on a T5 embedding of your text prompt. EnCodec decodes the predicted tokens back into audio.
Three consequences follow directly, and they explain nearly every rough edge you'll hit:
- Output is 32 kHz, mono for the base models. That is below CD rate — fine for sketching, noticeably band-limited next to a commercial master.
- The context window is 30 seconds (1,503 tokens). Anything longer is stitched together, and the model is not aware of what it wrote more than 30 seconds ago.
- Generation is sequential. Time scales linearly with duration — there is no way to parallelise within one clip. Batching several clips at once is free-ish; making one clip longer is not.
Two ways to run it: audiocraft or transformers
This is the first real decision, and most tutorials never mention it.
| audiocraft (Meta) | transformers (Hugging Face) | |
|---|---|---|
| Install | pip install audiocraft | pip install transformers torch |
| Dependency pins | Pins torch==2.1.0, documented for Python 3.9 — routinely fights modern environments | Tracks current torch and Python |
| Melody conditioning | Yes — generate_with_chroma | Via the separate MusicgenMelody classes |
| >30 s generation | Built in — automatic sliding window | Not built in; you roll it yourself |
| Loudness-normalised writing | Yes — audio_write | No, save it yourself |
| Maintenance | Effectively frozen | Actively maintained |
Pick audiocraft if you want melody conditioning or clips longer than 30 seconds without writing the stitching yourself. Pick transformers if you are putting MusicGen into an existing PyTorch project and don't want a pinned torch version dictating your whole dependency tree. The bulk of this guide uses audiocraft, with the transformers equivalent at the end.
Installing audiocraft without fighting it
AudioCraft's README asks for Python 3.9 and pins PyTorch 2.1.0. Installing it into an environment you care about is how people end up with a broken torch. Give it its own environment, always:
# a dedicated environment — do not install this next to anything you value
conda create -n musicgen python=3.9 -y
conda activate musicgen
# torch FIRST, matched to your CUDA version
python -m pip install 'torch==2.1.0'
python -m pip install setuptools wheel
python -m pip install -U audiocraft
Install torch before audiocraft. The package pulls in xformers, which compiles against whatever torch it finds, and letting pip resolve both at once is the single most common cause of install failure.
You also want ffmpeg on the system — audiocraft shells out to it for anything that isn't a WAV:
# macOS
brew install ffmpeg
# Debian/Ubuntu
sudo apt-get install ffmpeg
# or inside the conda env
conda install 'ffmpeg<5' -c conda-forge
Choosing a checkpoint
Six base models, plus stereo fine-tunes of each. The VRAM column is inference at fp32 with a short duration; long generations and large batches push it higher.
| Checkpoint | Params | VRAM (approx.) | Melody | Notes |
|---|---|---|---|---|
| facebook/musicgen-small | 300M | ~4 GB | No | The only one that's tolerable on CPU |
| facebook/musicgen-medium | 1.5B | ~16 GB | No | Meta's recommended quality/compute point |
| facebook/musicgen-large | 3.3B | ~24 GB | No | Best text-only quality |
| facebook/musicgen-melody | 1.5B | ~16 GB | Yes | Melody conditioning, the most interesting model here |
| facebook/musicgen-melody-large | 3.3B | ~24 GB | Yes | Melody at large scale |
| facebook/musicgen-stereo-* | as base | slightly higher | as base | Stereo fine-tunes of every model above |
Meta's own guidance is that medium or melody is the sweet spot, and that a GPU with at least 16 GB is needed for the 1.5B models. Below that, you are on small, or generating very short clips.
The stereo models are genuine fine-tunes, not a post-processing widener: they emit two interleaved EnCodec streams. If the output is going anywhere near a mix, start with musicgen-stereo-medium rather than widening mono afterwards.
Your first generation
from audiocraft.models import MusicGen
from audiocraft.data.audio import audio_write
model = MusicGen.get_pretrained('facebook/musicgen-small')
model.set_generation_params(duration=8)
wav = model.generate(['upbeat 80s synthwave with a driving bassline'])
for i, one in enumerate(wav):
audio_write(f'clip_{i}', one.cpu(), model.sample_rate,
strategy='loudness', loudness_compressor=True)
print(model.sample_rate) # 32000
print(model.audio_channels) # 1 for mono checkpoints
First run downloads several GB of weights to ~/.cache/huggingface. After that, load time is a few seconds.
audio_write adds the extension for you — pass 'clip_0', get clip_0.wav. Passing 'clip_0.wav' gets you clip_0.wav.wav, which catches everyone once.
Every generation parameter
These are the real defaults from MusicGen.set_generation_params. Most guides show you duration and stop.
| Parameter | Default | What it does |
|---|---|---|
| use_sampling | True | Sample from the distribution rather than taking the argmax. Setting this False gives greedy decoding, which sounds noticeably worse and repetitive. Leave it on. |
| top_k | 250 | Sample from the 250 most likely tokens. Lower (40–100) is more conservative and coherent; higher is more adventurous and more likely to wander. |
| top_p | 0.0 | Nucleus sampling. Zero means disabled, and top_k is used instead. Set it above 0 and it takes over. Try 0.9 if you want variety without the long tail. |
| temperature | 1.0 | Flattens or sharpens the distribution. Below 1.0 gives safer, more repetitive output; above 1.2 tends to disintegrate rhythmically. |
| duration | 30.0 | Seconds to generate. Above 30 triggers the sliding-window path below. |
| cfg_coef | 3.0 | Classifier-free guidance strength — how hard the model is pushed toward your prompt. This is the parameter most worth tuning. Raise to 4–5 to make it obey a complicated prompt; the cost is a thinner, more brittle sound. Below 2 it largely ignores you. |
| two_step_cfg | False | Runs conditional and unconditional passes separately instead of batched. Halves throughput, saves memory. A last resort when you're out of VRAM. |
| extend_stride | 18 | For generations past 30 s: how far the window advances each step. See below. Must be less than 30. |
In practice, cfg_coef and top_k are the two you'll actually reach for. A reasonable starting point for something that needs to follow a detailed prompt:
model.set_generation_params(
duration=20,
cfg_coef=4.5, # push harder toward the prompt
top_k=100, # tighter sampling, fewer odd detours
temperature=0.95, # marginally more conservative
)
Generating longer than 30 seconds
Ask for more than 30 seconds and audiocraft generates in overlapping windows: produce 30 s, advance by extend_stride, re-prime the model with the tail of what it just wrote, continue. The overlap is 30 − extend_stride seconds of context carried forward.
model.set_generation_params(
duration=90, # three windows
extend_stride=15, # 15 s of fresh audio per window, 15 s of overlap
)
wav = model.generate(['slow cinematic strings building to a climax'])
The trade-off is direct:
- Smaller stride (say 10) — 20 seconds of overlapping context, smoother joins, but many more windows and much slower generation.
- Larger stride (the default 18, or up to 29) — 12 seconds of context, faster, with a greater chance of an audible seam or a drift in key or tempo.
Melody conditioning
The melody models take a chroma representation of audio you supply and follow its harmonic contour while taking style from the text prompt. You hum the tune; it decides what the band sounds like.
import torchaudio
from audiocraft.models import MusicGen
from audiocraft.data.audio import audio_write
model = MusicGen.get_pretrained('facebook/musicgen-melody')
model.set_generation_params(duration=12, cfg_coef=3.0)
melody, sr = torchaudio.load('hummed_idea.wav')
wav = model.generate_with_chroma(
descriptions=['lofi hip hop, warm rhodes, vinyl crackle, brushed drums'],
melody_wavs=melody[None], # note the batch dimension
melody_sample_rate=sr,
progress=True,
)
audio_write('with_melody', wav[0].cpu(), model.sample_rate, strategy='loudness')
Notes that matter in practice:
- The input is resampled and converted internally — you don't need to match 32 kHz yourself.
- Conditioning uses chroma (pitch-class energy over time), not a note transcription. It tracks harmony, not melody exactly, and it ignores your timbre entirely.
- Only the first duration seconds of the reference are used.
- A clean monophonic input works far better than a dense mix. If your source is a full track, run it through Demucs and condition on a single stem.
- melody_wavs expects a batch dimension — melody[None]. Forgetting it is the most common error here.
Continuing existing audio
Different from melody conditioning: generate_continuation takes real audio as a prompt and keeps playing from where it stopped, in the same material.
import torchaudio
from audiocraft.models import MusicGen
from audiocraft.data.audio import audio_write
model = MusicGen.get_pretrained('facebook/musicgen-medium')
model.set_generation_params(duration=20)
prompt, sr = torchaudio.load('four_bar_loop.wav')
prompt = prompt[..., : int(5 * sr)] # first 5 seconds only
wav = model.generate_continuation(
prompt=prompt[None],
prompt_sample_rate=sr,
descriptions=['keep the groove, add a saxophone solo'],
progress=True,
)
audio_write('continued', wav[0].cpu(), model.sample_rate, strategy='loudness')
The prompt audio counts against the 30-second context. Feed it 20 seconds with duration=30 and you get 10 seconds of new music, not 30. Two to five seconds of prompt is usually enough to establish key and feel while leaving room to generate.
Batching, seeds and picking winners
MusicGen output is inconsistent enough that generating one clip and judging the model on it is a mistake. Generate several and audition. Multiple descriptions in one call are batched through the GPU together, which is far faster than looping:
prompts = [
'dark techno, rolling 909 hats, sub bass, 130 bpm',
'dark techno, rolling 909 hats, sub bass, 130 bpm',
'dark techno, rolling 909 hats, sub bass, 130 bpm',
'dark techno, distorted kick, acid 303 line, 132 bpm',
]
model.set_generation_params(duration=12)
wav = model.generate(prompts, progress=True)
for i, one in enumerate(wav):
audio_write(f'take_{i:02d}', one.cpu(), model.sample_rate, strategy='loudness')
Pass the same string several times to get variations on one prompt. For reproducibility, seed torch before generating — audiocraft has no seed argument of its own:
import torch
torch.manual_seed(42)
wav = model.generate(['ambient drone, deep and slow'])
Writing files properly
audio_write does more than dump a tensor, and its defaults are not what you want for a set of clips you plan to compare:
| Parameter | Default | What it does |
|---|---|---|
| format | 'wav' | Also 'mp3', 'ogg', 'flac' — the non-WAV formats need ffmpeg |
| strategy | 'peak' | 'peak', 'rms', 'loudness' or 'clip'. Use 'loudness' so clips are comparable by ear instead of by peak |
| loudness_headroom_db | 14 | Target headroom for the loudness strategy |
| loudness_compressor | False | Soft-limits before normalising; worth enabling for peaky material |
| mp3_rate | 320 | kbps, when format is mp3 |
| add_suffix | True | Appends the extension — the source of the double-.wav problem |
audio_write(
'render',
wav[0].cpu(),
model.sample_rate,
format='wav',
strategy='loudness',
loudness_headroom_db=14,
loudness_compressor=True,
add_suffix=True, # writes render.wav
)
Judging generations by peak normalisation is misleading — a quiet, dense mix and a loud sparse one hit the same peak. Loudness normalisation makes an A/B actually meaningful.
Progress callbacks
Long generations look like a hang. Wire up a callback so they don't:
def on_progress(done, total):
print(f'\r{done}/{total} tokens ({100 * done / total:.0f}%)', end='', flush=True)
model.set_custom_progress_callback(on_progress)
wav = model.generate(['epic orchestral trailer music'], progress=True)
The transformers route
If a pinned torch is unacceptable, the Hugging Face implementation avoids audiocraft entirely. Duration is expressed in tokens rather than seconds — at 50 tokens per second, max_new_tokens=256 is about 5.1 seconds:
import torch, scipy.io.wavfile
from transformers import AutoProcessor, MusicgenForConditionalGeneration
processor = AutoProcessor.from_pretrained('facebook/musicgen-small')
model = MusicgenForConditionalGeneration.from_pretrained(
'facebook/musicgen-small', device_map='auto')
inputs = processor(
text=['80s pop track with bassy drums and synth',
'90s rock song with loud guitars and heavy drums'],
padding=True,
return_tensors='pt',
)
audio_values = model.generate(**inputs, do_sample=True,
guidance_scale=3, max_new_tokens=512) # ~10 s
sr = model.config.audio_encoder.sampling_rate
scipy.io.wavfile.write('hf_out.wav', rate=sr, data=audio_values[0, 0].cpu().numpy())
guidance_scale here is the same concept as cfg_coef in audiocraft, with the same default of 3. Note there is no loudness normalisation and no built-in path past 30 seconds — you get a raw tensor and the rest is yours.
Writing prompts that work
MusicGen was trained on text-audio pairs from stock music libraries, which shapes what it responds to. Prompts that read like a stock-library description outperform poetic ones.
- Name instruments explicitly. "Rhodes electric piano, brushed drums, upright bass" beats "jazzy".
- State tempo in words and numbers. "Slow, 70 BPM" — it responds to the feel more than the figure, but both help.
- Include production language. "Warm analog tape saturation", "dry close-miked", "wide reverb" all land.
- Don't ask for vocals. Vocals were deliberately stripped from the training data using Demucs. Ask for singing and you get a wordless, uncanny approximation. For vocals you need ACE-Step or YuE.
- Don't name artists. Partly it doesn't work, and partly it's the kind of prompt that makes the output legally awkward.
- Don't stack contradictions. "Aggressive and calm, minimal but layered" produces mush. Raising cfg_coef on a contradictory prompt makes it worse, not better.
More on this in the prompt guide.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| CUDA out of memory | Drop to a smaller checkpoint, shorten duration, reduce batch size, or set two_step_cfg=True. Memory scales with all three of model size, duration and batch. |
| Cannot stride by more than max generation duration | extend_stride is ≥ 30. Set it lower. |
| xformers build fails on install | torch wasn't installed first, or your Python is too new for the pinned version. Fresh environment, torch first. |
| Output is silence or noise | Usually a mangled melody tensor — check the shape is (batch, channels, samples) and that melody[None] is there. |
| NotImplementedError on MPS | An op has no Metal kernel. Set PYTORCH_ENABLE_MPS_FALLBACK=1, or just use CPU. |
| ffmpeg not found when writing mp3 | Install ffmpeg system-wide, or write WAV and convert afterwards. |
| Every clip sounds identical | A seed is pinned somewhere, or use_sampling got set to False. |
Licensing: the part that matters commercially
The training data was licensed — Meta's own collection, Shutterstock and Pond5, about 20,000 hours — so the provenance is cleaner than some models. That doesn't loosen the weights licence.
If you need permissive terms, the practical open alternatives are ACE-Step (Apache-2.0) and Stable Audio Open (Stability AI Community License, free below a revenue threshold). Neither matches MusicGen on instrumental quality, and both are usable.
Where MusicGen actually fits
After all that, an honest summary. MusicGen is very good at eight to thirty seconds of instrumental texture in a recognisable genre, and at following a melody you give it. It is not good at song structure, vocals, long-form coherence, or clean high-end. Treat it as a sketch generator and a sample source — generate a lot, keep the 5% that's usable, and build around it in a DAW.
For that last part, pedalboard handles post-processing in the same Python script, and librosa will tell you what tempo and key you actually got, which is rarely what you asked for.
How this model’s licence compares with every other open model → · GPU & VRAM requirements →