๐Ÿฆ™Stalecollected in 35m

Qwen 3.5 Tops Non-Thinking Benchmarks

Qwen 3.5 Tops Non-Thinking Benchmarks
PostLinkedIn
๐Ÿฆ™Read original on Reddit r/LocalLLaMA
#benchmarks#open-source#non-thinking#moeqwen-3.5qwen-3.5glm-5minimax-m2deepseek-r1aa

๐Ÿ’กQwen 3.5-27B beats reasoning models like Deepseek R1 in speed benchmarks (37 pts)

โšก 30-Second TL;DR

What Changed

397B achieves 40 points on AA Intelligence Index, best open-source except GLM-5 (41)

Why It Matters

Showcases open-source LLMs closing gap on efficiency without reasoning overhead, ideal for low-latency apps. Boosts adoption of smaller Qwen variants for practical deployments.

What To Do Next

Download Qwen 3.5-27B from Hugging Face and benchmark on AA non-thinking tasks.

Who should care:Developers & AI Engineers

Key Points

  • โ€ข397B achieves 40 points on AA Intelligence Index, best open-source except GLM-5 (41)
  • โ€ข27B scores 37, surpasses Minimax M2 (36) and equals 35B-A3B thinking performance
  • โ€ข110B at 36 points, above gpt-oss-120b; 35B-A3B at 31, tops Deepseek R1 (19-27)
  • โ€ขEmphasizes compact efficiency in 27B dense model

๐Ÿง  Deep Insight

Background and context from public sources โ€” not the original article. 5 sources cited.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขQwen 3.5 series includes multimodal capabilities, processing text, images, and video inputs with text outputs across all four models: Flash, 35B-A3B, 122B-A10B, and 27B[1].
  • โ€ขModels are released under Apache 2.0 license, available on Hugging Face and ModelScope, with Qwen3.5-Flash offering 1M token context and API pricing at $0.10/M input and $0.40/M output tokens[1].
  • โ€ขQwen3.5 Small series (0.8B-9B parameters) emphasizes edge deployment with native multimodal architecture in the 4B model for superior spatial reasoning and OCR over adapter systems[3].
  • โ€ขQwen3.5-397B-A17B was released in mid-February 2026, marking the series start, with smaller models outperforming predecessors like Qwen3-235B-A22B due to improved architecture and RL[1].
๐Ÿ“Š Competitor Analysisโ–ธ Show
ModelKey FeaturesPricingBenchmarks
Qwen3.5-397BMultimodal (text/image/video), Apache 2.0, 1M context (Flash)API: $0.10/M in, $0.40/M outAA Intelligence 40 (2nd to GLM-5 at 41), tops non-thinking[1][5]
GLM-5Closed detailsN/AAA Intelligence 41[1]
GPT-5 miniProprietary multimodalClosedCompetitive target, lower cost claim[1]
Claude Sonnet 4.5ProprietaryClosedCompetitive target, lower cost claim[1]
Minimax M2Open details sparseN/AAA Intelligence 36 (below Qwen3.5-27B at 37)[1]

๐Ÿ› ๏ธ Technical Deep Dive

  • โ€ขQwen3.5-397B-A17B is a Mixture-of-Experts (MoE) model with 397B total parameters and 17B active parameters, enabling high efficiency[1][4].
  • โ€ขNative multimodal integration processes visual and textual tokens in a unified latent space from early training stages, improving spatial reasoning and OCR accuracy compared to adapter-based vision towers[1][3].
  • โ€ขQwen3.5 Small (0.8B-9B) uses Scaled Reinforcement Learning in the 9B model to optimize logical reasoning paths, rivaling models 5-10x larger; optimized for low VRAM and edge/IoT with ultra-low latency[3].
  • โ€ขSmaller models like 35B-A3B outperform larger predecessors (e.g., Qwen3-235B-A22B) via enhanced architecture, data quality, and RL, prioritizing compute efficiency[1].
  • โ€ขBenchmarks include BFCL-V4, VITA-Bench, DeepPlanning for overall performance ranking[5].

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Qwen 3.5 accelerates open-source multimodal agent adoption
Native multimodal architecture and Apache 2.0 licensing enable commercial edge deployments, closing gaps with proprietary models like GPT-5 mini at lower costs[1][3].
MoE efficiency shifts industry from scale to optimization
Smaller variants like 35B-A3B and 27B outperform much larger predecessors and rivals via RL and architecture, proving compute efficiency over raw parameters[1].
Edge AI proliferation via Qwen3.5 Small series
0.8B-9B models target consumer hardware and IoT with minimal VRAM and native multimodality, enabling privacy-focused local applications[3].

โณ Timeline

2026-02
Qwen3.5-397B-A17B released, initiating series with top non-thinking benchmark performance
2026-03
Qwen 3.5 lineup expanded to Flash, 35B-A3B, 122B-A10B, 27B, and Small series (0.8B-9B)
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.