Qwen3.5 vs Qwen3 Benchmarks Visualized

๐กVisual benchmark diffs show Qwen3.5 gains over Qwen3โpick winners fast.
โก 30-Second TL;DR
What Changed
Averaged scores from official release pages
Why It Matters
Provides quick performance comparison to guide model selection for research and deployment decisions.
What To Do Next
Access the Google Sheet linked in the post to analyze specific benchmark metrics.
Key Points
- โขAveraged scores from official release pages
- โขColor-coded bars: purple/blue for Qwen3.5, orange/yellow for Qwen3
- โขSorted by legend order, missing data for smaller models
- โขGoogle Sheet link for raw benchmark data
๐ง Deep Insight
Background and context from public sources โ not the original article. 8 sources cited.
๐ Enhanced Key Takeaways
- โขQwen3.5 models utilize Mixture-of-Experts (MoE) architecture, such as Qwen3.5-35B-A3B activating only 3B of 35B total parameters per token, enabling superior performance to larger predecessors at lower inference cost[1][3].
- โขQwen3.5-397B-A17B achieves 19x faster decoding on 256k token contexts compared to Qwen3-Max while matching reasoning and coding benchmarks[2].
- โขModels are released open-weight under Apache 2.0 license on platforms like Hugging Face, supporting unrestricted fine-tuning and deployment[1].
- โขQwen3.5 excels in multimodal tasks, scoring 90.8% on OmniDocBench v1.5 (outperforming GPT-5.2 and Claude Opus 4.5) and 67.5 on ERQA embodied reasoning[2].
๐ Competitor Analysisโธ Show
| Model | Architecture | Active Params | Key Benchmarks | Context Window | License |
|---|---|---|---|---|---|
| Qwen3.5-35B-A3B | MoE + Hybrid Attn | 3B | Surpasses Qwen3-235B-A22B; GPT-5-mini class | 262K | Apache 2.0 [1] |
| Qwen3.5-397B-A17B | Gated DeltaNet MoE | 17B | Matches Qwen3-Max reasoning; 90.8 OmniDocBench | 1M (Plus) | Apache 2.0 [2][3] |
| Claude Sonnet 4.6 | Dense | N/A | 6x slower than Qwen3.5-Plus | N/A | Proprietary [1] |
| GPT-5.2 | Dense | N/A | 85.7 OmniDocBench | N/A | Proprietary [2] |
๐ ๏ธ Technical Deep Dive
- โขQwen3.5-35B-A3B: 35B total parameters, 3B active (8.6% activation), MoE with hybrid attention, runs on 8GB+ VRAM GPUs via GGUF quantization[1].
- โขQwen3.5-397B-A17B: 397B total, 17B active, fuses Gated DeltaNet linear attention with high-sparsity MoE, supports FP8 precision and heterogeneous parallelism[3].
- โขQwen3.5-122B: Activates 10B of 122B parameters, leads in agentic benchmarks like BFCL-V4 (72.2), fits on NVIDIA DGX Spark[1].
- โขMultimodal capabilities via native training for image/video/document tasks, expanded to 201 languages with larger vocabulary[3].
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (8)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- digitalapplied.com โ Qwen 3 5 Medium Model Series Benchmarks Pricing Guide
- datacamp.com โ Qwen3 5
- slashdot.org โ Qwen vs Qwen3
- artificialanalysis.ai โ Qwen3 5 35b A3b vs Qwen3 5 27b
- news.ycombinator.com โ Item
- youtube.com โ Watch
- openrouter.ai โ Step 3.5 Flash
- ai-pricing.vercel.app โ Qwen Qwen3 5 Flash 02 23
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.

