⚛️較早收集於 2h

千問3.5霸榜全球開源大模型前四,10分鐘通過中級程式員5小時程式設計

PostLinkedIn
⚛️閱讀原文: 量子位
#benchmarks#coding-llm#downloadsqwen-3.5qwen-3.5

💡Open-source Qwen 3.5 crushes coding benchmarks: 10min vs 5hr pro task, 1B+ downloads

⚡ 30-Second TL;DR

有什麼變化

全球開源大模型排行前四

為什麼重要

此基準測試霸主地位提升開源AI採用率,挑戰封閉模型,並以優越程式設計效率加速開發者工作流程。

下一步行動

Download Qwen 3.5 from Hugging Face and test it on your intermediate coding benchmarks.

誰應關注:Developers & AI Engineers

關鍵要點

  • 全球開源大模型排行前四
  • 10分鐘完成中級程式員5小時程式任務
  • 累計下載量超10億
  • 衍生模型超過20萬個

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 6 個來源。

🔑 增強重點摘要

  • Qwen3.5-397B-A17B is a Mixture-of-Experts model with 397 billion total parameters and only 17 billion activated parameters per token[1][2][5].
  • Features a native context length of 262k tokens, enabled by Gated DeltaNet + Gated Attention hybrid mechanism[1][4].
  • Achieves top GPQA Diamond score of 88.4 among open-source models and excels in embodied reasoning with 67.5 on ERQA benchmark[2][3].
  • Released on February 16, 2026, under Apache 2.0 license as a native multimodal model processing text and images[4][5].
📊 競品分析▸ Show
Feature/BenchmarkQwen 3.5Kimi K2.5GLM-5Claude Opus 4.5
GPQA Diamond88.4[3]87.6[3]86.0[3]-
IFEval92.6[3]94.0[3]--
SWE-Bench (Verified/Pro)On par[1]-On par[1]Slightly above[1]
Parameters (Total/Active)397B/17B[5]1T/32B[3]-Closed
Context Length262k[1][4]262k[3]--

🛠️ 技術深入

  • Mixture-of-Experts (MoE) architecture: 397B total parameters, 17B active per token; 3x smaller than prior 235B-A22B but with 4x more experts plus a shared expert[1][5].
  • Attention mechanism: Gated DeltaNet + Gated Attention hybrid, supporting native 262k token context (vs. 32k/131k in prior models)[1].
  • Vocabulary size: 250k tokens with multi-token prediction, reducing costs by 10-60% across 201 languages[2].
  • Multimodal capabilities: Native text+image processing with early fusion for video; excels in document recognition (90.8% OmniDocBench)[2][5].
  • Efficiency: 19x faster decoding on 256k contexts than Qwen3-Max; 8.6x faster for standard tasks without performance loss[2].

🔮 前景展望AI analysis grounded in cited sources

Qwen3.5 will accelerate open-source adoption in agentic coding agents
Its SWE-Bench performance matches closed models like Claude Opus 4.5 while being efficient for local deployment at 51GB RAM[1].
Efficiency gains will lower inference costs for multimodal apps by 10-60%
Multi-token prediction and large vocabulary reduce token usage across 201 languages, combined with MoE sparsity[2].
Qwen3.5 sets new open-weight benchmark for long-context reasoning
262k native context and top GPQA Diamond score enable advanced RAG and multi-step tasks previously limited to closed models[1][3].

時間線

2026-02-16
Qwen3.5 series released, including 397B-A17B MoE model
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 量子位

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。