🏠Stalecollected in 3h

Stability AI Launches Stability Audio 3.0 Model Family

Stability AI Launches Stability Audio 3.0 Model Family
PostLinkedIn
🏠Read original on IT之家

💡New state-of-the-art audio model capable of 6-minute generation with a focus on legal data compliance.

⚡ 30-Second TL;DR

What Changed

Supports generation of full-length music tracks up to 6 minutes and 20 seconds.

Why It Matters

This release significantly improves the coherence and length of AI-generated music, making it a viable tool for professional creators. The focus on licensed data sets sets a new benchmark for legal compliance in generative audio.

What To Do Next

Download the open-source small or medium Stability Audio 3.0 weights from their repository to test local audio generation performance.

Who should care:Developers & AI Engineers

Key Points

  • Supports generation of full-length music tracks up to 6 minutes and 20 seconds.
  • Includes four model sizes, with smaller variants optimized for local device execution.
  • Trained on legally licensed datasets to ensure compliance with music industry standards.
  • Small and medium models are open-source, while the large model is API-exclusive.

🧠 Deep Insight

Web-grounded analysis with 15 cited sources.

🔑 Enhanced Key Takeaways

  • Stability Audio 3.0 introduces a new semantic-acoustic autoencoder architecture, enabling more flexible and longer audio generation with variable length control at per-second granularity.
  • The model family includes specific parameter counts: Small and Small SFX models have 459 million parameters, the Medium model has 1.4 billion parameters, and the Large model has 2.7 billion parameters.
  • The Small model is uniquely capable of full music composition on-device, allowing offline generation of complete musical tracks up to two minutes without short sample limits, a significant improvement over previous on-device versions.
  • Stability AI provides LoRA training documentation for the Small and Medium models, enabling users to fine-tune them on their own audio libraries, with enterprise customers receiving guided fine-tuning support.
  • The models incorporate inpainting features, allowing users to edit individual segments, modify multiple sections, or extend existing tracks beyond their original length through causal continuation.
📊 Competitor Analysis▸ Show

AI Music Generation Models Comparison (as of May 2026)

Feature / ModelStability Audio 3.0Suno AI (v5.5)Udio (v1.5)Google DeepMind Lyria 3 ProEleven Music
Max Track LengthUp to 6 minutes 20 seconds (Medium/Large); 2 minutes (Small/SFX)Up to 4+ minutesNot explicitly stated, but supports segment-based extensionsNot explicitly stated, API-grade modelNot explicitly stated
Open-Source WeightsSmall, Small SFX, Medium models are open-weightsNoNoNoNo
Licensing/Commercial UseTrained on legally licensed datasets; outputs owned by user (Community License); Enterprise License with legal indemnification for >$1M revenueStrong commercial potential (check paid plans)Not Copyright Cleared (as of Jan 2026 for Udio generally)Focus on commercially safer workflowsCreator-friendly licensing; Copyright Cleared
On-Device CapabilitySmall model enables full music composition on-device, offlineNoNoNoNo
Vocal SupportNot primary focus, but generates musicExceptional vocal synthesis and lyric accuracyHigh-fidelity instrumentals; detailed section-by-section editingAPI-grade music modelTop-tier voices
Editing FeaturesInpainting (edit segments, modify sections, extend tracks); LoRA fine-tuningStudio features (stem separation, extensions)Detailed section-by-section editing, inpainting, stem exportsNot explicitly detailed for end-usersBuilt-in audio editing, style exclusion controls
Pricing (approx.)API access for Large model; free for open-weightsFree tier with credits; Pro plans start ~$10/monthNot explicitly stated, but often compared to Suno$0.10 per 30 seconds (via fal.ai)$0.80 per output audio minute (via fal.ai)
Inference Time0.44s for 2 min (Small); 1.31s for 6:20 min (Medium) on H200 GPUFast generation (30-60 seconds for full song)Not explicitly statedNot explicitly statedNot explicitly stated

🛠️ Technical Deep Dive

  • Architecture: Stable Audio 3.0 utilizes a novel semantic-acoustic autoencoder that projects audio into a compact latent space, enabling efficient diffusion-based generation while preserving audio fidelity and encouraging semantic structure.
  • Model Components: The models are latent diffusion models comprising a variational autoencoder (VAE), a text encoder, and a U-Net-based conditioned diffusion model.
  • Variational Autoencoder (VAE): Compresses stereo audio into a data-compressed, noise-resistant, and invertible lossy latent encoding, facilitating faster generation and training compared to raw audio samples. It uses a fully-convolutional architecture based on the Descript Audio Codec for arbitrary-length audio encoding and decoding.
  • Text Encoder: Integrates text prompts using a pre-trained Contrastive Language Audio Pretraining (CLAP) model, built on a curated dataset to ensure strong connections between text features and corresponding sounds.
  • Diffusion Model: The U-Net architecture (with 907 million parameters in the original Stable Audio) incorporates residual layers, self-attention layers, and cross-attention layers to denoise input data based on text and time embeddings.
  • Parameter Counts: Stable Audio 3.0 Small SFX and Small models each have 459 million parameters. The Medium model has 1.4 billion parameters, and the Large model has 2.7 billion parameters.
  • Inference Speed: The Small and Small SFX models can produce two-minute tracks in 0.44 seconds on an H200 GPU. The Medium model generates tracks up to 6 minutes 20 seconds in 1.31 seconds on an H200 GPU. Generally, models can generate music and sounds in less than 2 seconds on an H200 GPU and a few seconds on a MacBook Pro M4.
  • Training Data: The models are trained on a combination of over 800,000 licensed audio files (music, sound effects, single-instrument stems) from the production library AudioSparx (totaling over 19,500 hours) and filtered Creative Commons recordings from Freesound.
  • Post-training: Adversarial post-training is applied to accelerate inference and enhance generation quality, reducing inference steps while improving fidelity and prompt adherence.

🔮 Future ImplicationsAI analysis grounded in cited sources

Professional music production will see increased adoption of AI tools for full-length track creation.
The ability of Stable Audio 3.0 to generate professional-grade music up to 6 minutes and 20 seconds with maintained musical structure, combined with legally licensed training data and legal indemnification for enterprise users, directly addresses key barriers for commercial integration in the music industry.
Mobile and consumer-grade devices will become viable platforms for comprehensive music composition.
The Stable Audio 3.0 Small model's capability for full, offline music composition on smartphones and consumer laptops democratizes access to advanced AI music generation, fostering innovation in mobile audio creation.
The AI music generation market will intensify its focus on transparent and legally compliant data sourcing.
Stability AI's explicit emphasis on licensed training data and legal indemnification for enterprise customers sets a precedent, pressuring competitors, especially those facing copyright lawsuits, to adopt similar legally safer workflows.

Timeline

2019
Stability AI founded
2023-09
Initial public release of Stable Audio, capable of generating 95 seconds of audio, trained on AudioSparx data.
2024-04
Stable Audio 2.0 released, increasing maximum track length to three minutes.
2024-06
Stable Audio Open, an open-source variant for shorter samples (up to 47 seconds), was released.
2025-05
Stability AI partnered with Arm to release Stable Audio Open Small, optimized for smartphones (up to 11 seconds).
2026-05
Stability Audio 3.0 model family launched, supporting up to 6 minutes 20 seconds, with open-weights for smaller models and on-device composition capabilities.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: IT之家