Stability AI Launches Stability Audio 3.0 Model Family

💡New state-of-the-art audio model capable of 6-minute generation with a focus on legal data compliance.
⚡ 30-Second TL;DR
What Changed
Supports generation of full-length music tracks up to 6 minutes and 20 seconds.
Why It Matters
This release significantly improves the coherence and length of AI-generated music, making it a viable tool for professional creators. The focus on licensed data sets sets a new benchmark for legal compliance in generative audio.
What To Do Next
Download the open-source small or medium Stability Audio 3.0 weights from their repository to test local audio generation performance.
Key Points
- •Supports generation of full-length music tracks up to 6 minutes and 20 seconds.
- •Includes four model sizes, with smaller variants optimized for local device execution.
- •Trained on legally licensed datasets to ensure compliance with music industry standards.
- •Small and medium models are open-source, while the large model is API-exclusive.
🧠 Deep Insight
Web-grounded analysis with 15 cited sources.
🔑 Enhanced Key Takeaways
- •Stability Audio 3.0 introduces a new semantic-acoustic autoencoder architecture, enabling more flexible and longer audio generation with variable length control at per-second granularity.
- •The model family includes specific parameter counts: Small and Small SFX models have 459 million parameters, the Medium model has 1.4 billion parameters, and the Large model has 2.7 billion parameters.
- •The Small model is uniquely capable of full music composition on-device, allowing offline generation of complete musical tracks up to two minutes without short sample limits, a significant improvement over previous on-device versions.
- •Stability AI provides LoRA training documentation for the Small and Medium models, enabling users to fine-tune them on their own audio libraries, with enterprise customers receiving guided fine-tuning support.
- •The models incorporate inpainting features, allowing users to edit individual segments, modify multiple sections, or extend existing tracks beyond their original length through causal continuation.
📊 Competitor Analysis▸ Show
AI Music Generation Models Comparison (as of May 2026)
| Feature / Model | Stability Audio 3.0 | Suno AI (v5.5) | Udio (v1.5) | Google DeepMind Lyria 3 Pro | Eleven Music |
|---|---|---|---|---|---|
| Max Track Length | Up to 6 minutes 20 seconds (Medium/Large); 2 minutes (Small/SFX) | Up to 4+ minutes | Not explicitly stated, but supports segment-based extensions | Not explicitly stated, API-grade model | Not explicitly stated |
| Open-Source Weights | Small, Small SFX, Medium models are open-weights | No | No | No | No |
| Licensing/Commercial Use | Trained on legally licensed datasets; outputs owned by user (Community License); Enterprise License with legal indemnification for >$1M revenue | Strong commercial potential (check paid plans) | Not Copyright Cleared (as of Jan 2026 for Udio generally) | Focus on commercially safer workflows | Creator-friendly licensing; Copyright Cleared |
| On-Device Capability | Small model enables full music composition on-device, offline | No | No | No | No |
| Vocal Support | Not primary focus, but generates music | Exceptional vocal synthesis and lyric accuracy | High-fidelity instrumentals; detailed section-by-section editing | API-grade music model | Top-tier voices |
| Editing Features | Inpainting (edit segments, modify sections, extend tracks); LoRA fine-tuning | Studio features (stem separation, extensions) | Detailed section-by-section editing, inpainting, stem exports | Not explicitly detailed for end-users | Built-in audio editing, style exclusion controls |
| Pricing (approx.) | API access for Large model; free for open-weights | Free tier with credits; Pro plans start ~$10/month | Not explicitly stated, but often compared to Suno | $0.10 per 30 seconds (via fal.ai) | $0.80 per output audio minute (via fal.ai) |
| Inference Time | 0.44s for 2 min (Small); 1.31s for 6:20 min (Medium) on H200 GPU | Fast generation (30-60 seconds for full song) | Not explicitly stated | Not explicitly stated | Not explicitly stated |
🛠️ Technical Deep Dive
- Architecture: Stable Audio 3.0 utilizes a novel semantic-acoustic autoencoder that projects audio into a compact latent space, enabling efficient diffusion-based generation while preserving audio fidelity and encouraging semantic structure.
- Model Components: The models are latent diffusion models comprising a variational autoencoder (VAE), a text encoder, and a U-Net-based conditioned diffusion model.
- Variational Autoencoder (VAE): Compresses stereo audio into a data-compressed, noise-resistant, and invertible lossy latent encoding, facilitating faster generation and training compared to raw audio samples. It uses a fully-convolutional architecture based on the Descript Audio Codec for arbitrary-length audio encoding and decoding.
- Text Encoder: Integrates text prompts using a pre-trained Contrastive Language Audio Pretraining (CLAP) model, built on a curated dataset to ensure strong connections between text features and corresponding sounds.
- Diffusion Model: The U-Net architecture (with 907 million parameters in the original Stable Audio) incorporates residual layers, self-attention layers, and cross-attention layers to denoise input data based on text and time embeddings.
- Parameter Counts: Stable Audio 3.0 Small SFX and Small models each have 459 million parameters. The Medium model has 1.4 billion parameters, and the Large model has 2.7 billion parameters.
- Inference Speed: The Small and Small SFX models can produce two-minute tracks in 0.44 seconds on an H200 GPU. The Medium model generates tracks up to 6 minutes 20 seconds in 1.31 seconds on an H200 GPU. Generally, models can generate music and sounds in less than 2 seconds on an H200 GPU and a few seconds on a MacBook Pro M4.
- Training Data: The models are trained on a combination of over 800,000 licensed audio files (music, sound effects, single-instrument stems) from the production library AudioSparx (totaling over 19,500 hours) and filtered Creative Commons recordings from Freesound.
- Post-training: Adversarial post-training is applied to accelerate inference and enhance generation quality, reducing inference steps while improving fidelity and prompt adherence.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (15)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: IT之家 ↗
