🦙Freshcollected in 78m

Qwen3.8-2.4T-A95B Launches

Qwen3.8-2.4T-A95B Launches
PostLinkedIn
🦙Read original on Reddit r/LocalLLaMA

💡A newly announced Qwen model may expand the options for local LLM evaluation.

⚡ 30-Second TL;DR

What Changed

The announced model is Qwen3.8-2.4T-A95B.

Why It Matters

A new Qwen model could be relevant to teams evaluating open or locally deployable LLMs. Its practical value cannot be assessed until the model card, license, and hardware requirements are confirmed.

What To Do Next

Check the official Qwen3.8-2.4T-A95B model card and run a small inference test before considering deployment.

Who should care:Developers & AI Engineers

Key Points

  • The announced model is Qwen3.8-2.4T-A95B.
  • The release was surfaced on Reddit’s r/LocalLLaMA community.
  • No technical specifications, licensing information, or benchmark results are included.

🧠 Deep Insight

AI-generated analysis for this event.

🔑 Enhanced Key Takeaways

  • Qwen3.8-2.4T-A95B is the first 'Max-class' model from Alibaba to receive an open-weight release, following its initial debut as a closed-source API-only model in July 2026.
  • The model utilizes a Mixture-of-Experts (MoE) architecture with 2.4 trillion total parameters and approximately 95 billion active parameters per query.
  • While the open-weight version is available for local deployment, it lacks certain features found in the official Qwen3.8-Max API, such as native vision input and the default 1-million-token context length.
  • The release is part of a broader rollout that includes a smaller, more accessible 27-billion-parameter dense model (Qwen3.8-27B) designed for consumer-grade hardware.
  • Licensing for the open-weight model includes specific caveats, allowing free use for internal purposes or for entities with less than $50 million in annual revenue, with restrictions for larger-scale commercial services.
📊 Competitor Analysis▸ Show
FeatureQwen3.8-2.4T-A95BFable 5Kimi K3
Architecture2.4T MoE (95B Active)Proprietary2.8T Open Model
AccessibilityOpen WeightsClosed APIOpen Weights
Primary StrengthCoding/Agentic TasksGeneral ReasoningLong-context/Agentic
PricingFree (with revenue caps)Premium APIVaries

🛠️ Technical Deep Dive

  • Architecture: Mixture-of-Experts (MoE) with 2.4T total parameters and 95B active parameters.
  • Hidden Dimension: 8192.
  • Layers: 92.
  • Token Embedding: 248,320 (padded).
  • Hidden Layout: 23 x (3 x (Gated DeltaNet -> MoE) -> 1 x (Gated Attention -> MoE)).
  • Inference: Requires significant VRAM; weights are provided in bf16 and fp8 formats, with community-driven GGUF quantizations emerging for local use.

🔮 Future ImplicationsAI analysis grounded in cited sources

Qwen3.8-27B will become the primary benchmark for local LLM performance in late 2026.
The 27B model is specifically optimized for consumer hardware, making it more accessible for widespread developer adoption compared to the 2.4T flagship.
Alibaba will shift toward a 'staggered' release strategy for future flagship models.
The successful staggered rollout of the Max-class model followed by the 27B variant demonstrates a strategy to maximize PR impact and community engagement.

Timeline

2026-07
Alibaba previews Qwen3.8-Max at the World AI Conference in Shanghai.
2026-08-03
Official release of Qwen3.8-Max via Alibaba Cloud API.
2026-08-12
Open weights for Qwen3.8-2.4T-A95B released on Hugging Face and ModelScope.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA