眾智FlagOS推出Qwen3.5 397B多芯版本

💡Largest open MoE VLM now runs out-of-box on NVIDIA & Chinese chips (397B params).
⚡ 30-Second TL;DR
有什麼變化
Qwen3.5-397B完整適配沐曦、平頭哥真武及NVIDIA晶片,精度對齊
為什麼重要
此次發布解決多晶片適配的核心痛點,讓開發者無需改碼即可在多元硬體上運行頂級開源MoE多模態模型,無性能損失。促進AI部署硬體多樣化,尤其有利中國生態。
下一步行動
Download FlagRelease/Qwen3.5-397B-A17B-nvidia-FlagOS from HuggingFace and deploy via vLLM-plugin-FL README.
關鍵要點
- •Qwen3.5-397B完整適配沐曦、平頭哥真武及NVIDIA晶片,精度對齊
- •vLLM-plugin-FL實現零改碼多晶片推理
- •BF16版本支援雙機16卡部署
- •FlagRelease於HuggingFace及魔搭提供直接下載
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 5 個來源。
🔑 增強重點摘要
- •FlagOS adapted Alibaba's Qwen3.5-397B-A17B, the largest open-source multimodal MoE model with 397B total parameters and 17B active parameters, for deployment on Metax, Zhenwu, and NVIDIA chips.[article]
- •vLLM-plugin-FL provides unified multi-chip inference with zero code changes, leveraging vLLM's high-throughput features like paged attention and continuous batching.[article][3]
- •Verified BF16 dual-machine 16-card setups enable seamless deployment, aligning full precision across diverse hardware.[article]
- •Ready-to-use models available on HuggingFace and ModelScope via FlagRelease.[article]
- •Qwen3.5-397B-A17B added to LMSYS Arena for Text, Vision, and Code benchmarks alongside Claude Sonnet 4.6.[5]
📊 競品分析▸ Show
| Feature | FlagOS Qwen3.5-397B-A17B | AirLLM | vLLM |
|---|---|---|---|
| Model Size | 397B total / 17B active MoE | Up to 405B | N/A (Inference Engine) |
| Hardware | Metax, Zhenwu, NVIDIA multi-chip | Low-memory GPUs (4-8GB) | Multi-GPU/node |
| Key Tech | vLLM-plugin-FL, BF16 16-card | Layer-by-layer loading | Paged attention, continuous batching |
| Benchmarks | LMSYS Arena (Text/Vision/Code) | Up to 3x speed w/ compression | High-throughput serving |
| Pricing | Open-source, free download | Open-source | Open-source |
🛠️ 技術深入
- Qwen3.5-397B-A17B is a multimodal Mixture-of-Experts (MoE) model with 397 billion total parameters but only 17 billion active per inference, optimizing compute efficiency.[article][5]
- Supports BF16 precision with full alignment for Metax (likely Chinese AI chip), Zhenwu, and NVIDIA GPUs, enabling dual-machine 16-card deployments.[article]
- vLLM-plugin-FL integrates with vLLM for zero-code-change inference, building on vLLM's paged attention, continuous batching, prefix caching, and multi-GPU support.[article][3]
🔮 前景展望AI analysis grounded in cited sources
FlagOS's multi-chip adaptation of Qwen3.5-397B-A17B democratizes access to massive open-source multimodal models across diverse hardware, reducing reliance on single-vendor ecosystems like NVIDIA and potentially accelerating adoption in cost-sensitive regions with chips like Metax and Zhenwu.
📎 來源 (5)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 机器之心 ↗
每週 AI 簡報
每週一封,可隨時退訂。