librosa vs Essentia
Both extract features from audio. librosa is the easier one to learn and prototype with; Essentia is faster, broader and built for production.
Side by side
| librosa | Essentia | |
|---|---|---|
| Core | Python + NumPy | C++ with Python bindings |
| Speed | Adequate | Much faster |
| Install | pip install librosa | Heavier |
| Pretrained models | No | Yes — genre, mood, instruments |
| Licence | ISC — permissive | AGPL-3.0 |
| Learning curve | Gentle | Steeper |
| Language | Python | Python |
| License | ISC | AGPL-3.0-only |
| Platforms | macOS Windows Linux | macOS Windows Linux iOS Android |
| First released | 2013 | 2013 |
| Maintained | Yes | Yes |
Where they actually differ
AGPL is the first thing to check
Essentia is AGPL-3.0, which reaches across a network boundary — offering a service built on it can oblige you to release your source. librosa is ISC, which is permissive and unproblematic. For commercial work this can decide the matter before any technical comparison.
Essentia is dramatically faster
A C++ core makes a real difference on large catalogues. For a handful of files nobody notices; for a hundred thousand, librosa becomes the bottleneck and Essentia does not.
Pretrained models are Essentia’s other advantage
Genre, mood, danceability and instrument classifiers ship ready to use. Getting equivalent high-level descriptors from librosa means training something yourself on top of its features.
librosa is much easier to live in
Clean NumPy-native API, excellent documentation, and it behaves exactly as expected in a notebook. For exploration, teaching and papers it remains the default for good reason.
Which should you choose?
Choose librosa when…
- You are prototyping, researching or teaching
- You want a permissive licence
- You work in notebooks with NumPy
- You need clear documentation and examples
Choose Essentia when…
- You are processing a large catalogue
- You need genre, mood or instrument classification
- Speed is a real constraint
- AGPL is acceptable for your distribution