🤖較早收集於 37h

397B Qwen3-Next 達 1T 效能—如何達成?

PostLinkedIn
🤖閱讀原文: Reddit r/MachineLearning
#scaling#synthetic-data#inference-speedqwen3-next

💡397B model at 1T perf? Arch or data tricks—key for efficient LLM scaling

⚡ 30-Second TL;DR

有什麼變化

397B 模型達 1T 效能指標

為什麼重要

突顯高效大型模型推論潛在突破,相關於高吞吐 LLM 部署。

下一步行動

Review Qwen3-Next benchmarks on Hugging Face to compare 397B inference speeds.

誰應關注:Researchers & Academics

關鍵要點

  • 397B 模型達 1T 效能指標
  • 辯論:Qwen3-Next 純架構獲益或合成資料蒸餾
  • 引發大型 LLM 擴展效率疑問

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 6 個來源。

🔑 增強重點摘要

  • Qwen3.5-397B-A17B uses a Hybrid Mixture-of-Experts (MoE) architecture with 397 billion total parameters but only 17 billion active per token, enabling 1T-level performance through efficiency gains[1][2].
  • The model achieves 19x faster decoding on long-context tasks (256k tokens) and 8.6x faster for standard workflows compared to Qwen3-Max, while matching its reasoning and coding capabilities[1].
  • FP8 precision reduces memory usage by 50% and boosts speeds by over 10% at trillion-token scale, combined with high-quality visual-text data filtering to rival larger 1T-parameter models[1].
  • Features native multimodality with early fusion vision-language training, supporting chat, RAG, vision-language understanding, video understanding, and agentic workflows[2][5].
  • Positioned as competitive with top models like Gemini 3 Pro and Claude Opus, with strong benchmark performance but not claiming SOTA in coding[3][4].
📊 競品分析▸ Show
Feature/BenchmarkQwen3.5-397B-A17BQwen3-MaxQwen3-Next-80B-A3B
Total Parameters397B (17B active)>1T80B (3B active)
Speed (vs Qwen3-Max)19x faster (long-context)BaselineN/A
BenchmarksMatches reasoning/coding; outperforms Qwen3-VLStrong baselineOutperformed in 10 benchmarks (e.g., GPQA, LiveCodeBench)[1][3]
Context Length262k native (up to 1M)N/AN/A
PricingCost-efficient (50% less memory)HigherN/A

🛠️ 技術深入

  • Architecture: Hybrid MoE with 512 total experts (10 routed + 1 shared per token); 60 layers; hidden dimension 4,096; layout: 15 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE))[2].
  • Attention: Gated DeltaNet (64 linear heads for V, 16 for QK, head dim 128); Gated Attention (32 heads for Q, 2 for KV, head dim 256; RoPE dim 64)[2].
  • MoE Details: Expert intermediate dimension 1,024; vocabulary 248,320; input context 262,144 tokens (extensible to 1,010,000 via YaRN)[2].
  • Multimodal: Early fusion vision-language training; supports text/video inputs; operates in thinking mode with reasoning details[2][6].
  • Optimizations: FP8 pipeline for 50% memory reduction; NVIDIA GPU-optimized for faster inference[1][2].

🔮 前景展望AI analysis grounded in cited sources

Qwen3.5-397B-A17B demonstrates MoE efficiency can deliver 1T-scale performance from sub-400B models, lowering costs and enabling broader deployment of multimodal agents; accelerates Chinese open model competition, pressuring labs like DeepSeek for v4 refresh while advancing native spatial intelligence and agentic workflows[1][4].

時間線

2026-02
Qwen3.5-397B-A17B released as efficient MoE multimodal model matching Qwen3-Max performance with 19x speed gains[1][2][4]
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Reddit r/MachineLearning

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。