🧠較早收集於 5m

眾智FlagOS推出Qwen3.5 397B多芯版本

眾智FlagOS推出Qwen3.5 397B多芯版本
PostLinkedIn
🧠閱讀原文: 机器之心
#moe-model#multi-chip#bf16flagos

💡Largest open MoE VLM now runs out-of-box on NVIDIA & Chinese chips (397B params).

⚡ 30-Second TL;DR

有什麼變化

Qwen3.5-397B完整適配沐曦、平頭哥真武及NVIDIA晶片,精度對齊

為什麼重要

此次發布解決多晶片適配的核心痛點,讓開發者無需改碼即可在多元硬體上運行頂級開源MoE多模態模型,無性能損失。促進AI部署硬體多樣化,尤其有利中國生態。

下一步行動

Download FlagRelease/Qwen3.5-397B-A17B-nvidia-FlagOS from HuggingFace and deploy via vLLM-plugin-FL README.

誰應關注:Developers & AI Engineers

關鍵要點

  • Qwen3.5-397B完整適配沐曦、平頭哥真武及NVIDIA晶片,精度對齊
  • vLLM-plugin-FL實現零改碼多晶片推理
  • BF16版本支援雙機16卡部署
  • FlagRelease於HuggingFace及魔搭提供直接下載

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 5 個來源。

🔑 增強重點摘要

  • FlagOS adapted Alibaba's Qwen3.5-397B-A17B, the largest open-source multimodal MoE model with 397B total parameters and 17B active parameters, for deployment on Metax, Zhenwu, and NVIDIA chips.[article]
  • vLLM-plugin-FL provides unified multi-chip inference with zero code changes, leveraging vLLM's high-throughput features like paged attention and continuous batching.[article][3]
  • Verified BF16 dual-machine 16-card setups enable seamless deployment, aligning full precision across diverse hardware.[article]
  • Ready-to-use models available on HuggingFace and ModelScope via FlagRelease.[article]
  • Qwen3.5-397B-A17B added to LMSYS Arena for Text, Vision, and Code benchmarks alongside Claude Sonnet 4.6.[5]
📊 競品分析▸ Show
FeatureFlagOS Qwen3.5-397B-A17BAirLLMvLLM
Model Size397B total / 17B active MoEUp to 405BN/A (Inference Engine)
HardwareMetax, Zhenwu, NVIDIA multi-chipLow-memory GPUs (4-8GB)Multi-GPU/node
Key TechvLLM-plugin-FL, BF16 16-cardLayer-by-layer loadingPaged attention, continuous batching
BenchmarksLMSYS Arena (Text/Vision/Code)Up to 3x speed w/ compressionHigh-throughput serving
PricingOpen-source, free downloadOpen-sourceOpen-source

🛠️ 技術深入

  • Qwen3.5-397B-A17B is a multimodal Mixture-of-Experts (MoE) model with 397 billion total parameters but only 17 billion active per inference, optimizing compute efficiency.[article][5]
  • Supports BF16 precision with full alignment for Metax (likely Chinese AI chip), Zhenwu, and NVIDIA GPUs, enabling dual-machine 16-card deployments.[article]
  • vLLM-plugin-FL integrates with vLLM for zero-code-change inference, building on vLLM's paged attention, continuous batching, prefix caching, and multi-GPU support.[article][3]

🔮 前景展望AI analysis grounded in cited sources

FlagOS's multi-chip adaptation of Qwen3.5-397B-A17B democratizes access to massive open-source multimodal models across diverse hardware, reducing reliance on single-vendor ecosystems like NVIDIA and potentially accelerating adoption in cost-sensitive regions with chips like Metax and Zhenwu.

📎 來源 (5)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. westurner.github.io — Hnlog
  2. news.ycombinator.com — Item
  3. ludwigabap.com — Bookmarks
  4. t.me — Githubtrending
  5. xagi.in — AI News
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 机器之心

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。