Run Stable Audio Open Locally
Stereo 44.1 kHz audio from a text prompt, on your own GPU, under a licence you can actually ship with. The complete reference for running Stability AI’s open audio model in Python.
Last updated 2026-09-12
Stable Audio Open is the open-weights text-to-audio model from Stability AI. It generates up to 47 seconds of stereo 44.1 kHz audio — a real sample rate, unlike most open models — and its licence permits commercial use below a revenue threshold. That combination makes it the most practical open model for anyone shipping a product.
It is also, by Stability's own admission, better at sound effects than at music. Knowing that up front saves a lot of disappointment.
What it is and what it isn't
Three components: an autoencoder (Oobleck) that compresses waveforms into latents, a T5 text encoder for conditioning, and a transformer diffusion model (DiT) operating in latent space. It is a diffusion model, not an autoregressive one — which is why it takes a step count rather than generating left to right, and why it can produce the entire clip at once.
| Property | Value |
|---|---|
| Output | 44.1 kHz stereo |
| Maximum length | 47 seconds (the default audio_end_in_s is 47.55) |
| Training data | ~48k recordings, ~47k from Freesound plus Free Music Archive — all CC0, CC BY or CC Sampling+ |
| Licence | Stability AI Community License |
| Access | Gated on Hugging Face — accept terms before the weights download |
| Vocals | No. The model card is explicit: it cannot generate realistic vocals |
| Languages | English prompts only |
That training set is the interesting part. Everything in it is Creative Commons, which is why Stability can offer commercial terms at all, and also why the model is strongest on the kind of material Freesound is full of: field recordings, impacts, textures, foley, one-shots, loops. It is weaker on full musical arrangements than MusicGen, which was trained on stock music.
Two routes: diffusers or stable-audio-tools
Both run the same weights.
| diffusers | stable-audio-tools | |
|---|---|---|
| Install | pip install diffusers transformers | Clone or pip install stable-audio-tools |
| Best for | Inference, batching, putting it in an app | Training, fine-tuning, the Gradio UI |
| API stability | Stable, well documented | Research-oriented, moves faster |
| Quantisation | Supported via bitsandbytes | Not out of the box |
Use diffusers unless you intend to fine-tune. That's the route this guide follows, with the stable-audio-tools path covered near the end.
Getting access to the weights
The repository is gated. Before any code runs you have to accept the licence on the model page, or the download fails with a 401 that doesn't explain itself.
- Sign in at Hugging Face and open stabilityai/stable-audio-open-1.0.
- Accept the licence and share contact info as prompted.
- Create an access token, then authenticate locally.
pip install -U huggingface_hub
hf auth login # paste your token when prompted
Install
python -m venv .venv && source .venv/bin/activate
pip install torch --index-url https://download.pytorch.org/whl/cu121
pip install diffusers transformers accelerate soundfile
soundfile is what writes the WAV; torch should be a build that matches your CUDA version. No pinned-version minefield here — this is a much better-behaved install than audiocraft.
Your first generation
import torch
import soundfile as sf
from diffusers import StableAudioPipeline
pipe = StableAudioPipeline.from_pretrained(
'stabilityai/stable-audio-open-1.0',
torch_dtype=torch.float16,
)
pipe = pipe.to('cuda')
generator = torch.Generator('cuda').manual_seed(0)
audio = pipe(
'The sound of a hammer hitting a wooden surface.',
negative_prompt='Low quality.',
num_inference_steps=200,
audio_end_in_s=10.0,
num_waveforms_per_prompt=3,
generator=generator,
).audios
# (channels, samples) -> (samples, channels) for soundfile
output = audio[0].T.float().cpu().numpy()
sf.write('hammer.wav', output, pipe.vae.sampling_rate)
Two details in that snippet that are easy to get wrong and produce a broken file:
- audio[0].T — the pipeline returns (channels, samples) and soundfile wants (samples, channels). Skip the transpose and you get a two-sample file or an error.
- pipe.vae.sampling_rate — read the rate from the model rather than hardcoding 44100. It's correct by construction and it survives model changes.
Every parameter that matters
| Parameter | Default | What it does |
|---|---|---|
| prompt | — | Your description. Descriptive and specific beats short and genre-y. |
| negative_prompt | None | What to steer away from. Unusually effective here — the documented recommendation of "low quality, average quality" is a genuine, free quality improvement. Use it by default. |
| num_inference_steps | 100 | Denoising steps. The main quality/time dial. 50 for drafts, 100 for normal use, 200 when it's a final render. Returns diminish sharply past 200. |
| audio_end_in_s | 47.55 | Clip length in seconds. Shorter is faster and usually tighter — this model wanders when given 47 seconds. |
| audio_start_in_s | 0.0 | Start offset. Conditioning signal, not a trim — the model was trained with timing information and uses it. |
| guidance_scale | 7.0 | Prompt adherence. Higher follows the text more literally with more artefacts; 5–9 is the useful band. |
| num_waveforms_per_prompt | 1 | Generate several and have the pipeline rank them against the prompt automatically, best first. The cheapest quality win available. |
| generator | None | A seeded torch.Generator for reproducibility. |
| initial_audio_waveforms | None | Audio-to-audio: start from existing audio instead of noise. See below. |
| output_type | 'pt' | 'pt' for a tensor, 'np' for numpy, 'latent' for the raw latent. |
The negative prompt is not optional
On most diffusion models a negative prompt is a nice-to-have. Here it's the difference between usable and muddy, and Stability documents the specific recommendation:
audio = pipe(
'Warm analog synth pad, slow attack, gentle movement, high quality',
negative_prompt='low quality, average quality, distortion, clipping',
num_inference_steps=150,
audio_end_in_s=20.0,
).audios
Beyond the generic quality terms, negative prompts are the best tool for removing things the model likes to add unprompted. Asking for a clean drum loop and getting room reverb? Add "reverb, room, ambience" to the negative. Getting hiss on a synth pad? "noise, hiss, distortion".
Generate several, keep the best
num_waveforms_per_prompt does something more useful than looping: it generates a batch and scores every candidate against the prompt text, returning them ranked best to worst. Index zero is the model's own pick.
audio = pipe(
'128 BPM techno drum loop, punchy kick, crisp hats, dry, high quality',
negative_prompt='low quality, average quality',
num_inference_steps=100,
audio_end_in_s=8.0,
num_waveforms_per_prompt=5, # ranked best-first
).audios
for i, wav in enumerate(audio):
sf.write(f'loop_rank{i}.wav', wav.T.float().cpu().numpy(), pipe.vae.sampling_rate)
This is a much better use of compute than one 200-step generation. Five candidates at 100 steps will nearly always beat one at 500.
Seeds and reproducibility
Diffusion makes genuine A/B testing possible in a way autoregressive models don't — fix the seed and change one thing:
seed = 1234
for steps in (50, 100, 200):
g = torch.Generator('cuda').manual_seed(seed)
audio = pipe('rain on a tin roof, close, high quality',
negative_prompt='low quality, average quality',
num_inference_steps=steps,
audio_end_in_s=10.0,
generator=g).audios
sf.write(f'rain_{steps}steps.wav',
audio[0].T.float().cpu().numpy(), pipe.vae.sampling_rate)
Note the device string in torch.Generator("cuda") must match where the pipeline lives, or the seed is silently ignored and you'll wonder why nothing is reproducible.
Audio-to-audio: start from your own sound
Rather than starting from pure noise, seed the diffusion with existing audio. This is how you get variations on a sample you already have, or push a rough recording toward a described sound.
import torchaudio
waveform, sr = torchaudio.load('my_loop.wav')
audio = pipe(
'the same rhythm played on wooden percussion, dry, high quality',
negative_prompt='low quality, average quality',
num_inference_steps=100,
audio_end_in_s=10.0,
initial_audio_waveforms=waveform.unsqueeze(0).to('cuda'),
initial_audio_sampling_rate=sr,
).audios
The initial audio's sample rate must be declared, and is resampled internally. Effect strength isn't a single dial — reduce num_inference_steps to stay closer to the source, raise it to travel further toward the prompt.
Running on less VRAM
Full fp16 inference wants roughly 8 GB. Three levers, in order of what to try first:
# 1. fp16 instead of fp32 — roughly halves memory, negligible quality cost
pipe = StableAudioPipeline.from_pretrained(
'stabilityai/stable-audio-open-1.0', torch_dtype=torch.float16)
# 2. move components to GPU only while they run
pipe.enable_model_cpu_offload()
# 3. shorter clips — memory scales with audio_end_in_s
audio = pipe(prompt, audio_end_in_s=10.0).audios
If that isn't enough, 8-bit quantisation via bitsandbytes brings it down further at some cost in quality:
import torch
from diffusers import BitsAndBytesConfig as DiffusersBitsAndBytesConfig
from diffusers import StableAudioDiTModel, StableAudioPipeline
from transformers import BitsAndBytesConfig, T5EncoderModel
quant_config = BitsAndBytesConfig(load_in_8bit=True)
text_encoder_8bit = T5EncoderModel.from_pretrained(
'stabilityai/stable-audio-open-1.0', subfolder='text_encoder',
quantization_config=quant_config, dtype=torch.float16)
quant_config = DiffusersBitsAndBytesConfig(load_in_8bit=True)
transformer_8bit = StableAudioDiTModel.from_pretrained(
'stabilityai/stable-audio-open-1.0', subfolder='transformer',
quantization_config=quant_config, dtype=torch.float16)
pipe = StableAudioPipeline.from_pretrained(
'stabilityai/stable-audio-open-1.0',
text_encoder=text_encoder_8bit,
transformer=transformer_8bit,
dtype=torch.float16,
device_map='balanced',
)
Batch generation for sample packs
Where this model earns its keep: generating a hundred variations overnight and keeping the twenty that are usable.
import itertools, pathlib, torch, soundfile as sf
sources = ['analog synth', 'plucked string', 'metallic']
characters = ['short and percussive', 'long and evolving', 'gritty and distorted']
out = pathlib.Path('pack'); out.mkdir(exist_ok=True)
for i, (src, char) in enumerate(itertools.product(sources, characters)):
prompt = f'{src} one-shot, {char}, dry, high quality'
g = torch.Generator('cuda').manual_seed(i)
audio = pipe(prompt,
negative_prompt='low quality, average quality, reverb',
num_inference_steps=100,
audio_end_in_s=4.0,
num_waveforms_per_prompt=3,
generator=g).audios
for rank, wav in enumerate(audio):
name = f'{src}_{char}_{rank}'.replace(' ', '_')
sf.write(out / f'{name}.wav',
wav.T.float().cpu().numpy(), pipe.vae.sampling_rate)
Prompting Stable Audio Open
The training data was Freesound, and Freesound's descriptions are literal and technical. Prompt in that register.
- Describe the sound, not the vibe. "Metallic impact with long decay, recorded in a large hall" beats "ominous".
- Use quality adjectives. "High quality", "clear", "professional recording" measurably help — they correlate with better source material in the training set.
- Name the recording context. "Close-miked", "field recording", "studio", "mono cassette" all land.
- Give BPM for loops. "128 BPM techno drum loop" produces something loopable far more often than "techno drums".
- Be specific over short. Stability's own tip: "melodic techno with a fast beat and synths" outperforms "techno".
- Don't ask for singing, lyrics or named artists. The first two don't work at all; the third mostly doesn't and creates problems if it does.
The stable-audio-tools route
Stability's own repository, aimed at training and fine-tuning rather than inference. Development targets Python 3.10 and it requires PyTorch 2.5 or later for Flash Attention support — note that's a very different torch requirement from audiocraft's, which is a good argument for separate environments per model.
pip install stable-audio-tools
# Gradio UI against the released weights
python3 ./run_gradio.py --pretrained-name stabilityai/stable-audio-open-1.0
# fine-tuning on your own dataset
python3 ./train.py --dataset-config /path/to/dataset.json \
--model-config /path/to/model.json \
--name my_finetune
Reach for this when you want to fine-tune on your own sample library — the repo carries the training scripts, dataset configs and an unwrapping step to turn a training checkpoint into a usable model. For plain generation, diffusers is less work.
Stable Audio Open Small
In May 2025 Stability and Arm released a 341M-parameter variant designed to run on Arm CPUs — around 11 seconds of audio in under 8 seconds on a phone, no GPU and no cloud. It carries the same Community License and works with stable-audio-tools.
It is meaningfully lower quality than the full model and limited to short clips. But for on-device sound generation in a game or an app, it is currently the only realistic open option.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| 401 / gated repo error on download | Licence not accepted, or not logged in. Accept on the model page, then hf auth login. |
| WAV is garbage or almost empty | Missing the .T transpose before sf.write. |
| CUDA out of memory | fp16, then enable_model_cpu_offload(), then shorten the clip, then quantise. |
| Output is muddy or noisy | Add the negative prompt, raise steps to 150–200, raise num_waveforms_per_prompt. |
| Seed doesn't reproduce | Generator device doesn't match the pipeline device. |
| Clip drifts or falls apart near the end | Normal at long durations. Generate shorter and concatenate. |
| Asked for vocals, got noise | Working as documented — the model cannot do vocals. |
Licensing
Stable Audio Open is released under the Stability AI Community License. In broad terms that permits research use and commercial use by organisations below an annual revenue threshold, with separate commercial terms required above it. The threshold and the exact conditions are Stability's to set and have changed before — read the current licence at stability.ai/license before you ship, rather than trusting any summary, including this one.
The training data being entirely Creative Commons is a real advantage over models trained on unclear provenance, and it is why these terms are possible at all.
Where it fits
Stable Audio Open is the open model to reach for when you need audio rather than songs, at a sample rate you can put in a mix, under terms you can actually ship. Sound design, foley, transitions, textures, loops, sample-pack generation.
For full songs with vocals, go to ACE-Step. For instrumental musical passages with melody conditioning, MusicGen is still stronger. For processing what comes out, pedalboard is in the same Python process.
How this model’s licence compares with every other open model → · GPU & VRAM requirements →