Qwen 3.5 Tops Non-Thinking Benchmarks

๐กQwen 3.5-27B beats reasoning models like Deepseek R1 in speed benchmarks (37 pts)
โก 30-Second TL;DR
What Changed
397B achieves 40 points on AA Intelligence Index, best open-source except GLM-5 (41)
Why It Matters
Showcases open-source LLMs closing gap on efficiency without reasoning overhead, ideal for low-latency apps. Boosts adoption of smaller Qwen variants for practical deployments.
What To Do Next
Download Qwen 3.5-27B from Hugging Face and benchmark on AA non-thinking tasks.
Key Points
- โข397B achieves 40 points on AA Intelligence Index, best open-source except GLM-5 (41)
- โข27B scores 37, surpasses Minimax M2 (36) and equals 35B-A3B thinking performance
- โข110B at 36 points, above gpt-oss-120b; 35B-A3B at 31, tops Deepseek R1 (19-27)
- โขEmphasizes compact efficiency in 27B dense model
๐ง Deep Insight
Background and context from public sources โ not the original article. 5 sources cited.
๐ Enhanced Key Takeaways
- โขQwen 3.5 series includes multimodal capabilities, processing text, images, and video inputs with text outputs across all four models: Flash, 35B-A3B, 122B-A10B, and 27B[1].
- โขModels are released under Apache 2.0 license, available on Hugging Face and ModelScope, with Qwen3.5-Flash offering 1M token context and API pricing at $0.10/M input and $0.40/M output tokens[1].
- โขQwen3.5 Small series (0.8B-9B parameters) emphasizes edge deployment with native multimodal architecture in the 4B model for superior spatial reasoning and OCR over adapter systems[3].
- โขQwen3.5-397B-A17B was released in mid-February 2026, marking the series start, with smaller models outperforming predecessors like Qwen3-235B-A22B due to improved architecture and RL[1].
๐ Competitor Analysisโธ Show
| Model | Key Features | Pricing | Benchmarks |
|---|---|---|---|
| Qwen3.5-397B | Multimodal (text/image/video), Apache 2.0, 1M context (Flash) | API: $0.10/M in, $0.40/M out | AA Intelligence 40 (2nd to GLM-5 at 41), tops non-thinking[1][5] |
| GLM-5 | Closed details | N/A | AA Intelligence 41[1] |
| GPT-5 mini | Proprietary multimodal | Closed | Competitive target, lower cost claim[1] |
| Claude Sonnet 4.5 | Proprietary | Closed | Competitive target, lower cost claim[1] |
| Minimax M2 | Open details sparse | N/A | AA Intelligence 36 (below Qwen3.5-27B at 37)[1] |
๐ ๏ธ Technical Deep Dive
- โขQwen3.5-397B-A17B is a Mixture-of-Experts (MoE) model with 397B total parameters and 17B active parameters, enabling high efficiency[1][4].
- โขNative multimodal integration processes visual and textual tokens in a unified latent space from early training stages, improving spatial reasoning and OCR accuracy compared to adapter-based vision towers[1][3].
- โขQwen3.5 Small (0.8B-9B) uses Scaled Reinforcement Learning in the 9B model to optimize logical reasoning paths, rivaling models 5-10x larger; optimized for low VRAM and edge/IoT with ultra-low latency[3].
- โขSmaller models like 35B-A3B outperform larger predecessors (e.g., Qwen3-235B-A22B) via enhanced architecture, data quality, and RL, prioritizing compute efficiency[1].
- โขBenchmarks include BFCL-V4, VITA-Bench, DeepPlanning for overall performance ranking[5].
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (5)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- the-decoder.com โ Alibabas Open Qwen 3 5 Takes Aim at Gpt 5 Mini and Claude Sonnet 4 5 at a Fraction of the Cost
- binance.com โ 297427083420257
- marktechpost.com โ Alibaba Just Released Qwen 3 5 Small Models a Family of 0 8b to 9b Parameters Built for on Device Applications
- youtube.com โ Watch
- qwen.ai โ Blog
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.