💼較早收集於 18m

Qwen 3.5 397B 以低成本擊敗萬億參數模型

Qwen 3.5 397B 以低成本擊敗萬億參數模型
PostLinkedIn
💼閱讀原文: VentureBeat
#moe-architecture#multimodal#context-windowqwen3.5-397b-a17b

💡Open-weight MoE beats trillion-param model at 1/18th Gemini cost, 19x faster inference

⚡ 30-Second TL;DR

有什麼變化

3970 億參數,每 token 僅激活 170 億,透過 512 個 MoE 專家

為什麼重要

此發布挑戰企業 AI 採購,提供可部署、可擁有的開源權重模型,具旗艦性能,減少對租用萬億參數巨頭的依賴。它加速大規模高上下文、多模態 AI 採用,大幅降低生產工作負載的推論費用。

下一步行動

Download Qwen3.5-397B-A17B from Hugging Face and benchmark inference speed on your GPU cluster.

誰應關注:Enterprise & Security Teams

關鍵要點

  • 3970 億參數,每 token 僅激活 170 億,透過 512 個 MoE 專家
  • 256K 上下文快 19 倍於 Qwen3-Max,運行成本降 60%
  • 原生多模態訓練,涵蓋文字/圖像/影片
  • 成本僅 Gemini 3 Pro 的 1/18,處理 8 倍並發工作負載
  • 多 token 預測與優化注意力機制提升速度

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 6 個來源。

🔑 增強重點摘要

  • Qwen3.5-397B-A17B achieves 19x faster decoding on long-context tasks (256K tokens) compared to Qwen3-Max while matching its reasoning and coding performance[2]
  • The model uses a Hybrid Mixture-of-Experts architecture with 512 total experts, activating only 10 routed + 1 shared expert per token for efficiency[1]
  • Native multimodal training on trillions of tokens across text, image, and video domains covering 201 languages enables early fusion vision-language capabilities[1]
  • Qwen3.5-Plus hosted version extends context window to 1 million tokens with adaptive tool use (web search, code interpreters) via Alibaba Cloud Model Studio[5]
  • MMMLU benchmark score of 88.5 represents significant improvement over Qwen3-Max (84.4) but remains below Gemini 3 Pro (90.6)[2]
📊 競品分析▸ Show
FeatureQwen3.5-397B-A17BQwen3-MaxGemini 3 Pro
Total Parameters397B~1T (estimated)Undisclosed
Active Parameters17BN/AN/A
Native Context256K (extensible to 1M via YaRN)UndisclosedUndisclosed
MMMLU Score88.584.490.6
Decoding Speed (256K)Baseline (19x faster than Qwen3-Max)1x referenceN/A
MultimodalityNative (text, image, video)Text-primaryNative
ArchitectureHybrid MoE with Gated DeltaNet + AttentionStandardN/A

🛠️ 技術深入

Architecture: 60 layers with hidden dimension 4,096; layout alternates 15 blocks of (3× Gated DeltaNet→MoE) and (1× Gated Attention→MoE)[1]Attention Mechanism: Gated DeltaNet uses 64 linear attention heads for values, 16 for QK with 128-dim heads; Gated Attention uses 32 heads for Q, 2 for KV with 256-dim heads and RoPE dimension 64[1]Expert Configuration: 512 total experts with 1,024 intermediate dimension; 11 experts activated per token (10 routed + 1 shared)[1]Context Extension: Native 262,144 token input context length, extensible to 1,010,000 tokens via YaRN RoPE scaling[1]Vocabulary: 248,320 tokens[1]Memory Requirements: ~800GB VRAM for full FP16/BF16 model; ~220GB for 4-bit quantization; runnable on Mac Studio/Pro with M-series Ultra (256GB RAM) or multi-GPU clusters[2]Training Data: Trillions of multimodal tokens across image, text, and video; 201 languages and dialects; training labeling and collection automated[1]Inference Modes: Thinking mode (internal reasoning) and Fast mode for standard workflows[2]

🔮 前景展望AI analysis grounded in cited sources

Qwen3.5-397B-A17B represents a significant shift in the open-source LLM landscape by demonstrating that sparse mixture-of-experts architectures can match or exceed dense trillion-parameter models in reasoning and coding while dramatically reducing computational costs and inference latency. This challenges the prevailing assumption that scale alone determines capability, potentially accelerating adoption of efficient open-weight models in enterprise and research settings. The native multimodal training from inception—rather than bolted-on vision adapters—establishes a new standard for vision-language model design. The 1M token context window in the hosted Qwen3.5-Plus variant enables new use cases in document analysis, long-form reasoning, and agentic workflows. However, the MMMLU gap versus Gemini 3 Pro (88.5 vs 90.6) suggests frontier performance still favors proprietary models, though the cost and speed advantages may offset this for many applications. The model's availability as open-weight software could intensify competition in the API market and influence how other labs (DeepSeek, Minimax, Kimi) design their next-generation models.

時間線

2024-11
Qwen3-Max released as previous generation flagship model
2025-Q4
Qwen3-VL introduced with vision-language capabilities
2026-02
Qwen3.5-397B-A17B launched with native multimodal training and Hybrid MoE architecture

📎 來源 (6)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. build.nvidia.com — Modelcard
  2. datacamp.com — Qwen3 5
  3. latent.space — Ainews Qwen35 397b A17b the Smallest
  4. qwen.ai — Research
  5. qwen.ai — Blog
  6. openrouter.ai — Qwen3.5 397b A17b
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: VentureBeat

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。