The Patchbay_

Run Stable Audio Open Locally

Stereo 44.1 kHz audio from a text prompt, on your own GPU, under a licence you can actually ship with. The complete reference for running Stability AI’s open audio model in Python.

Last updated 2026-09-12

Stable Audio Open is the open-weights text-to-audio model from Stability AI. It generates up to 47 seconds of stereo 44.1 kHz audio — a real sample rate, unlike most open models — and its licence permits commercial use below a revenue threshold. That combination makes it the most practical open model for anyone shipping a product.

It is also, by Stability's own admission, better at sound effects than at music. Knowing that up front saves a lot of disappointment.

What it is and what it isn't

Three components: an autoencoder (Oobleck) that compresses waveforms into latents, a T5 text encoder for conditioning, and a transformer diffusion model (DiT) operating in latent space. It is a diffusion model, not an autoregressive one — which is why it takes a step count rather than generating left to right, and why it can produce the entire clip at once.

PropertyValue
Output44.1 kHz stereo
Maximum length47 seconds (the default audio_end_in_s is 47.55)
Training data~48k recordings, ~47k from Freesound plus Free Music Archive — all CC0, CC BY or CC Sampling+
LicenceStability AI Community License
AccessGated on Hugging Face — accept terms before the weights download
VocalsNo. The model card is explicit: it cannot generate realistic vocals
LanguagesEnglish prompts only

That training set is the interesting part. Everything in it is Creative Commons, which is why Stability can offer commercial terms at all, and also why the model is strongest on the kind of material Freesound is full of: field recordings, impacts, textures, foley, one-shots, loops. It is weaker on full musical arrangements than MusicGen, which was trained on stock music.

Use it for what it's good at. If you need a whooshing transition, a rain texture, a metallic impact, a four-bar analog loop or a drone bed, this is the best open model available. If you need a two-minute song with structure, it is the wrong tool and no amount of prompting will fix it.

Two routes: diffusers or stable-audio-tools

Both run the same weights.

 diffusersstable-audio-tools
Installpip install diffusers transformersClone or pip install stable-audio-tools
Best forInference, batching, putting it in an appTraining, fine-tuning, the Gradio UI
API stabilityStable, well documentedResearch-oriented, moves faster
QuantisationSupported via bitsandbytesNot out of the box

Use diffusers unless you intend to fine-tune. That's the route this guide follows, with the stable-audio-tools path covered near the end.

Getting access to the weights

The repository is gated. Before any code runs you have to accept the licence on the model page, or the download fails with a 401 that doesn't explain itself.

  1. Sign in at Hugging Face and open stabilityai/stable-audio-open-1.0.
  2. Accept the licence and share contact info as prompted.
  3. Create an access token, then authenticate locally.
pip install -U huggingface_hub
hf auth login          # paste your token when prompted

Install

python -m venv .venv && source .venv/bin/activate
pip install torch --index-url https://download.pytorch.org/whl/cu121
pip install diffusers transformers accelerate soundfile

soundfile is what writes the WAV; torch should be a build that matches your CUDA version. No pinned-version minefield here — this is a much better-behaved install than audiocraft.

Your first generation

import torch
import soundfile as sf
from diffusers import StableAudioPipeline

pipe = StableAudioPipeline.from_pretrained(
    'stabilityai/stable-audio-open-1.0',
    torch_dtype=torch.float16,
)
pipe = pipe.to('cuda')

generator = torch.Generator('cuda').manual_seed(0)

audio = pipe(
    'The sound of a hammer hitting a wooden surface.',
    negative_prompt='Low quality.',
    num_inference_steps=200,
    audio_end_in_s=10.0,
    num_waveforms_per_prompt=3,
    generator=generator,
).audios

# (channels, samples) -> (samples, channels) for soundfile
output = audio[0].T.float().cpu().numpy()
sf.write('hammer.wav', output, pipe.vae.sampling_rate)

Two details in that snippet that are easy to get wrong and produce a broken file:

Every parameter that matters

ParameterDefaultWhat it does
prompt—Your description. Descriptive and specific beats short and genre-y.
negative_promptNoneWhat to steer away from. Unusually effective here — the documented recommendation of "low quality, average quality" is a genuine, free quality improvement. Use it by default.
num_inference_steps100Denoising steps. The main quality/time dial. 50 for drafts, 100 for normal use, 200 when it's a final render. Returns diminish sharply past 200.
audio_end_in_s47.55Clip length in seconds. Shorter is faster and usually tighter — this model wanders when given 47 seconds.
audio_start_in_s0.0Start offset. Conditioning signal, not a trim — the model was trained with timing information and uses it.
guidance_scale7.0Prompt adherence. Higher follows the text more literally with more artefacts; 5–9 is the useful band.
num_waveforms_per_prompt1Generate several and have the pipeline rank them against the prompt automatically, best first. The cheapest quality win available.
generatorNoneA seeded torch.Generator for reproducibility.
initial_audio_waveformsNoneAudio-to-audio: start from existing audio instead of noise. See below.
output_type'pt''pt' for a tensor, 'np' for numpy, 'latent' for the raw latent.

The negative prompt is not optional

On most diffusion models a negative prompt is a nice-to-have. Here it's the difference between usable and muddy, and Stability documents the specific recommendation:

audio = pipe(
    'Warm analog synth pad, slow attack, gentle movement, high quality',
    negative_prompt='low quality, average quality, distortion, clipping',
    num_inference_steps=150,
    audio_end_in_s=20.0,
).audios

Beyond the generic quality terms, negative prompts are the best tool for removing things the model likes to add unprompted. Asking for a clean drum loop and getting room reverb? Add "reverb, room, ambience" to the negative. Getting hiss on a synth pad? "noise, hiss, distortion".

Generate several, keep the best

num_waveforms_per_prompt does something more useful than looping: it generates a batch and scores every candidate against the prompt text, returning them ranked best to worst. Index zero is the model's own pick.

audio = pipe(
    '128 BPM techno drum loop, punchy kick, crisp hats, dry, high quality',
    negative_prompt='low quality, average quality',
    num_inference_steps=100,
    audio_end_in_s=8.0,
    num_waveforms_per_prompt=5,   # ranked best-first
).audios

for i, wav in enumerate(audio):
    sf.write(f'loop_rank{i}.wav', wav.T.float().cpu().numpy(), pipe.vae.sampling_rate)

This is a much better use of compute than one 200-step generation. Five candidates at 100 steps will nearly always beat one at 500.

Seeds and reproducibility

Diffusion makes genuine A/B testing possible in a way autoregressive models don't — fix the seed and change one thing:

seed = 1234

for steps in (50, 100, 200):
    g = torch.Generator('cuda').manual_seed(seed)
    audio = pipe('rain on a tin roof, close, high quality',
                 negative_prompt='low quality, average quality',
                 num_inference_steps=steps,
                 audio_end_in_s=10.0,
                 generator=g).audios
    sf.write(f'rain_{steps}steps.wav',
             audio[0].T.float().cpu().numpy(), pipe.vae.sampling_rate)

Note the device string in torch.Generator("cuda") must match where the pipeline lives, or the seed is silently ignored and you'll wonder why nothing is reproducible.

Audio-to-audio: start from your own sound

Rather than starting from pure noise, seed the diffusion with existing audio. This is how you get variations on a sample you already have, or push a rough recording toward a described sound.

import torchaudio

waveform, sr = torchaudio.load('my_loop.wav')

audio = pipe(
    'the same rhythm played on wooden percussion, dry, high quality',
    negative_prompt='low quality, average quality',
    num_inference_steps=100,
    audio_end_in_s=10.0,
    initial_audio_waveforms=waveform.unsqueeze(0).to('cuda'),
    initial_audio_sampling_rate=sr,
).audios

The initial audio's sample rate must be declared, and is resampled internally. Effect strength isn't a single dial — reduce num_inference_steps to stay closer to the source, raise it to travel further toward the prompt.

Running on less VRAM

Full fp16 inference wants roughly 8 GB. Three levers, in order of what to try first:

# 1. fp16 instead of fp32 — roughly halves memory, negligible quality cost
pipe = StableAudioPipeline.from_pretrained(
    'stabilityai/stable-audio-open-1.0', torch_dtype=torch.float16)

# 2. move components to GPU only while they run
pipe.enable_model_cpu_offload()

# 3. shorter clips — memory scales with audio_end_in_s
audio = pipe(prompt, audio_end_in_s=10.0).audios

If that isn't enough, 8-bit quantisation via bitsandbytes brings it down further at some cost in quality:

import torch
from diffusers import BitsAndBytesConfig as DiffusersBitsAndBytesConfig
from diffusers import StableAudioDiTModel, StableAudioPipeline
from transformers import BitsAndBytesConfig, T5EncoderModel

quant_config = BitsAndBytesConfig(load_in_8bit=True)
text_encoder_8bit = T5EncoderModel.from_pretrained(
    'stabilityai/stable-audio-open-1.0', subfolder='text_encoder',
    quantization_config=quant_config, dtype=torch.float16)

quant_config = DiffusersBitsAndBytesConfig(load_in_8bit=True)
transformer_8bit = StableAudioDiTModel.from_pretrained(
    'stabilityai/stable-audio-open-1.0', subfolder='transformer',
    quantization_config=quant_config, dtype=torch.float16)

pipe = StableAudioPipeline.from_pretrained(
    'stabilityai/stable-audio-open-1.0',
    text_encoder=text_encoder_8bit,
    transformer=transformer_8bit,
    dtype=torch.float16,
    device_map='balanced',
)

Batch generation for sample packs

Where this model earns its keep: generating a hundred variations overnight and keeping the twenty that are usable.

import itertools, pathlib, torch, soundfile as sf

sources = ['analog synth', 'plucked string', 'metallic']
characters = ['short and percussive', 'long and evolving', 'gritty and distorted']

out = pathlib.Path('pack'); out.mkdir(exist_ok=True)

for i, (src, char) in enumerate(itertools.product(sources, characters)):
    prompt = f'{src} one-shot, {char}, dry, high quality'
    g = torch.Generator('cuda').manual_seed(i)
    audio = pipe(prompt,
                 negative_prompt='low quality, average quality, reverb',
                 num_inference_steps=100,
                 audio_end_in_s=4.0,
                 num_waveforms_per_prompt=3,
                 generator=g).audios
    for rank, wav in enumerate(audio):
        name = f'{src}_{char}_{rank}'.replace(' ', '_')
        sf.write(out / f'{name}.wav',
                 wav.T.float().cpu().numpy(), pipe.vae.sampling_rate)

Prompting Stable Audio Open

The training data was Freesound, and Freesound's descriptions are literal and technical. Prompt in that register.

The stable-audio-tools route

Stability's own repository, aimed at training and fine-tuning rather than inference. Development targets Python 3.10 and it requires PyTorch 2.5 or later for Flash Attention support — note that's a very different torch requirement from audiocraft's, which is a good argument for separate environments per model.

pip install stable-audio-tools

# Gradio UI against the released weights
python3 ./run_gradio.py --pretrained-name stabilityai/stable-audio-open-1.0

# fine-tuning on your own dataset
python3 ./train.py --dataset-config /path/to/dataset.json \
                   --model-config /path/to/model.json \
                   --name my_finetune

Reach for this when you want to fine-tune on your own sample library — the repo carries the training scripts, dataset configs and an unwrapping step to turn a training checkpoint into a usable model. For plain generation, diffusers is less work.

Stable Audio Open Small

In May 2025 Stability and Arm released a 341M-parameter variant designed to run on Arm CPUs — around 11 seconds of audio in under 8 seconds on a phone, no GPU and no cloud. It carries the same Community License and works with stable-audio-tools.

It is meaningfully lower quality than the full model and limited to short clips. But for on-device sound generation in a game or an app, it is currently the only realistic open option.

Troubleshooting

SymptomCause and fix
401 / gated repo error on downloadLicence not accepted, or not logged in. Accept on the model page, then hf auth login.
WAV is garbage or almost emptyMissing the .T transpose before sf.write.
CUDA out of memoryfp16, then enable_model_cpu_offload(), then shorten the clip, then quantise.
Output is muddy or noisyAdd the negative prompt, raise steps to 150–200, raise num_waveforms_per_prompt.
Seed doesn't reproduceGenerator device doesn't match the pipeline device.
Clip drifts or falls apart near the endNormal at long durations. Generate shorter and concatenate.
Asked for vocals, got noiseWorking as documented — the model cannot do vocals.

Licensing

Stable Audio Open is released under the Stability AI Community License. In broad terms that permits research use and commercial use by organisations below an annual revenue threshold, with separate commercial terms required above it. The threshold and the exact conditions are Stability's to set and have changed before — read the current licence at stability.ai/license before you ship, rather than trusting any summary, including this one.

The training data being entirely Creative Commons is a real advantage over models trained on unclear provenance, and it is why these terms are possible at all.

Compared with the alternatives. MusicGen's weights are CC-BY-NC — non-commercial, full stop. ACE-Step is Apache-2.0, the most permissive of the three. Stable Audio Open sits between them: commercially usable under conditions, with the cleanest training-data story.

Where it fits

Stable Audio Open is the open model to reach for when you need audio rather than songs, at a sample rate you can put in a mix, under terms you can actually ship. Sound design, foley, transitions, textures, loops, sample-pack generation.

For full songs with vocals, go to ACE-Step. For instrumental musical passages with melody conditioning, MusicGen is still stronger. For processing what comes out, pedalboard is in the same Python process.

Generate music with MusicGenStronger on musical passages, and it does melody conditioning.Generate full songs with ACE-StepApache-2.0, full songs with vocals, very fast.Stable Audio Open in the directorySpecs, alternatives and where it sits among the tools.

← AI music generation hub