13 Months of Local LLM Inference Leap

๐กLocal inference: $600 PC now runs Qwen3.5-35B at 20 tps โ huge progress!
โก 30-Second TL;DR
What Changed
$6000 DeepSeek R1 Q8 at 5 tps vs $600 PC Qwen3-27B Q4 at same speed
Why It Matters
Demonstrates explosive growth in affordable local AI, empowering practitioners with high-speed inference on budget hardware.
What To Do Next
Quantize Qwen3.5-35B-A3B to Q4 and benchmark 17-20 tps on a $600 mini PC.
Key Points
- โข$6000 DeepSeek R1 Q8 at 5 tps vs $600 PC Qwen3-27B Q4 at same speed
- โขQwen3.5-35B-A3B Q4/Q5 hits 17-20 tps on mini PC
- โขHighlights rapid gains in smaller model quality and local hardware
- โขPredicts 4B model outperforming Kimi 2.5 next year
๐ง Deep Insight
Background and context from public sources โ not the original article. 8 sources cited.
๐ Enhanced Key Takeaways
- โขQwen3 suite includes 8 models ranging from 6B to 235B parameters, with MoE variants like 235B (22B active) and 30B (3B active) enabling high performance at lower inference costs.[1][5]
- โขQwen3-235B outperforms DeepSeek R1 in ArenaHard (95.6 vs lower), AIMEโ24/โ25 math (85.7/81.4), and coding benchmarks, while Qwen3-4B achieves 76.6 ArenaHard and 73.8 AIMEโ24 despite small size.[1]
- โขQwen3 introduces hybrid thinking mode for complex reasoning toggled with non-thinking for speed, supports 119 languages, and uses diverse training data including web and PDFs, doubling prior datasets.[5][6]
๐ Competitor Analysisโธ Show
| Model | Key Benchmarks | Architecture | License |
|---|---|---|---|
| Qwen3-235B | ArenaHard: 95.6, AIMEโ24: 85.7 | MoE (235B total, 22B active) | Apache 2.0 |
| DeepSeek R1 | MATH-500: 97.3%, Codeforces: 96.3%, MMLU: 90.8% | Reasoning-first, 671B | MIT |
| Qwen3-30B-A3B | Strong coding/math vs GPT-4o, Gemma 3 | MoE (30B total, 3B active) | Apache 2.0 |
๐ ๏ธ Technical Deep Dive
- โขQwen3 MoE models reduce active parameters (e.g., 235B total with 22B active, 30B with 3B active) for inference efficiency comparable to dense models while expanding knowledge capacity 3x+.[1][5][6]
- โขHybrid thinking mode allows toggling between resource-intensive reasoning and fast non-thinking responses, with stable budget control improving math/coding/science accuracy.[1][5]
- โขDeepSeek R1 employs reasoning-first architecture with visible chain-of-thought, self-verification, and step-by-step debugging for transparent logical consistency.[3][6]
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (8)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- datacamp.com โ Qwen3
- composio.dev โ Qwen 3 vs Deepseek R1 Complete Comparison
- entelligence.ai โ Claude 4 vs Deepseek R1 vs Qwen 3
- composio.dev โ Qwq 32b vs Gemma 3 Mistral Small vs Deepseek R1
- youtube.com โ Watch
- digitalapplied.com โ Deepseek R1 vs Qwen 3 vs Mistral Large Comparison
- llm-stats.com โ Deepseek R1 Distill Qwen 32b vs Gemma 3 27b It
- youtube.com โ Watch
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
