35B A3B MoE Faster Than 9B Locally
๐กReal-world speeds: 35B MoE at 55t/s on 16GB VRAM beats slower 9B
โก 30-Second TL;DR
What Changed
35B A3B MoE runs at steady 49-55 tokens/s on 16GB VRAM and 64GB RAM
Why It Matters
Highlights practical local run speeds for MoE models on consumer hardware, guiding hardware choices for LLM inference.
What To Do Next
Test Qwen 3.5 35B-A3B MoE on your 16GB VRAM setup using llama.cpp.
Key Points
- โข35B A3B MoE runs at steady 49-55 tokens/s on 16GB VRAM and 64GB RAM
- โข9B variant slower at 23 t/s on same hardware
- โขUser praises 35B MoE as amazing for local inference
- โข9B rumored better than 120B OSS despite speed
๐ง Deep Insight
Background and context from public sources โ not the original article. 6 sources cited.
๐ Enhanced Key Takeaways
- โขQwen3.5-35B-A3B is a Mixture-of-Experts (MoE) model with 35B total parameters but only activates ~3B parameters per token via a routing system, enabling high speed on consumer hardware.[1]
- โขQwen3.5 series, including the 35B-A3B MoE, was officially released by Alibaba's Qwen team on February 24, 2026, and made available on Hugging Face.[5][6]
- โขThe 35B-A3B excels in speed (60-100+ t/s on RTX 3090/4090) but trails dense models like Qwen3.5-27B (15-25 t/s) in logic, coding, and complex reasoning due to fewer active parameters.[1]
๐ ๏ธ Technical Deep Dive
- โขQwen3.5-35B-A3B uses MoE architecture where a router activates ~3B parameters out of 35B total for each token, approximating the compute of a 3B dense model while retaining larger model knowledge.[1]
- โขFine-tuning Qwen3.5-35B-A3B requires 74GB VRAM for bf16 LoRA; Unsloth enables 1.5x faster training with 50% less VRAM, supports MoE kernels, and recommends bf16 over 4-bit QLoRA for stability.[3]
- โขCommunity estimates effective intelligence via โ(Total ร Active) formula, placing 35B-A3B ~10B dense equivalent, between 7B-14B in capability.[1]
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
๐ Sources (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA โ
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
