397B Qwen3.5 at 9 tok/s on $2100 desktop

💡Run 400B MoE at 9 tok/s on $2K PC – local AI inference breakthrough
⚡ 30-Second TL;DR
What Changed
Runs 397B Qwen3.5 at 5-9 tok/s on two $500 GPUs + 32GB RAM
Why It Matters
Revolutionizes local inference for massive MoE models, making flagship AI accessible on consumer desktops without cloud dependency. Lowers barriers for builders experimenting with 400B-scale models at home.
What To Do Next
Download Qwen3.5 Q4_K_M quants and test FOMOE on dual RTX 4090 setup.
Key Points
- •Runs 397B Qwen3.5 at 5-9 tok/s on two $500 GPUs + 32GB RAM
- •60% VRAM hit rate reduces NVMe reads to 28%, CAR drops to 7%
- •Dual GPU ping-pong overlaps loading and compute
- •3.5% perplexity drop on wikitext with Cache-Aware Routing
- •15K lines of C/HIP code for consumer hardware inference
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.