๐Ÿฆ™Stalecollected in 29m

Qwen 3 32B Beats All Qwen 3.5 in Evals

PostLinkedIn
๐Ÿฆ™Read original on Reddit r/LocalLLaMA
#benchmarks#moe#local-llmqwenqwen-3qwen-3.5

๐Ÿ’กDense 32B beats 397B MoEโ€”redefines local LLM efficiency picks.

โšก 30-Second TL;DR

What Changed

Qwen 3 32B (dense 32B) scores 9.63, beats Qwen 3.5 397B-A17B (9.40)

Why It Matters

Dense models challenge MoE hype; boosts local LLM options for consumer hardware.

What To Do Next

Run local benchmarks on Qwen 3 32B vs 35B-A3B for your hardware setup.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขQwen 3 32B (dense 32B) scores 9.63, beats Qwen 3.5 397B-A17B (9.40)
  • โ€ขQwen 3.5 35B-A3B wins 4 evals with 3B active params, 0.54 score/sec
  • โ€ขCoder Next ranks 7th, loses to generalists on coding tasks like SQL, Go
  • โ€ข58.5% valid judgments; top ranks stable despite noise

๐Ÿง  Deep Insight

Background and context from public sources โ€” not the original article. 8 sources cited.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขQwen 3 models support dual-mode operation with thinking mode for chain-of-thought reasoning on benchmarks like AIME25, achieving up to 92.3% accuracy, toggled via tokenizer flags for optimized latency.[1]
  • โ€ขQwen 3 excels in multilingual tasks across 119+ languages, with 70% accuracy on dialectal inference and strong performance in agentic tasks scoring 65.1 on Tau2-Bench.[1]
  • โ€ขQwen 3.5 offers API pricing at $0.40 per million input tokens and $1.20 output, with a 1 million token context window, multimodal support, and Apache 2.0 licensing for self-hosting.[3]
๐Ÿ“Š Competitor Analysisโ–ธ Show
BenchmarkQwen 3.5Gemini 3.1 ProClaude Opus 4.6GPT-5.3 CodexGrok 4.20
ARC-AGI-212%77.1%68.8%52.9%~16%
GPQA Diamond88.4%94.3%91.3%92.4%~88%
SWE-Bench76.4%80.6%80.8%โ€”~72โ€“75%
Pricing (input/output per 1M tokens)$0.40/$1.20Not specifiedNot specifiedNot specifiedNot specified

๐Ÿ› ๏ธ Technical Deep Dive

  • โ€ขQwen 3 variants feature hybrid reasoning mode: thinking mode activates step-by-step logic for math/coding, non-thinking for dialogue, supporting 32K contexts offline on edge devices.[1]
  • โ€ขSupports 119+ languages with nuanced instruction-following, refined on 36 trillion tokens including synthetic math/code data, ideal for federated learning on 16โ€“24GB VRAM with LoRA/QLoRA.[1]
  • โ€ขQwen 3.5 includes 1M token context window, multimodal inputs (text, image, audio), and Apache 2.0 open licensing for self-hosting.[3]

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Qwen 3 dense models like 32B will prioritize local inference over larger MoE variants
Blind evals show 32B topping MoE like 397B-A17B, while 35B-A3B suits hardware with 3B active params at 0.54 score/sec.[article]
Cost-sensitive deployments will favor Qwen 3.5 self-hosting
At $0.40/$1.20 per million tokens with Apache 2.0 licensing and 1M context, it undercuts Western models significantly.[3]
Multilingual edge AI will adopt Qwen 3 for IoT
119+ language support, 32K offline contexts, and Tau2-Bench agentic scores enable low-latency privacy-focused apps.[1]

โณ Timeline

2024-09
Qwen 2.5 release with strong coder and math variants
2025-01
Qwen 2.5-Max and 72B models benchmarked on MMLU, HumanEval
2026-01
Qwen 3 series launched with multilingual and MoE variants
2026-02
Qwen 3.5 introduced with 1M context and multimodal features
2026-03
Qwen 3 32B evals show it beating Qwen 3.5 models in blind tests
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ†—

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.