
397B Qwen3.5 at 9 tok/s on $2100 desktop
FOMOE enables running the 397B Qwen3.5 MoE model at 5-9 tokens/s on a $2,100 desktop with two $500 GPUs and 32GB RAM using Q4_K_M quants. It uses VRAM caching for common experts, rolling caches, and Cache-Aware Routing (CAR) to slash NVMe reads to 7%. Achieves this with dual-GPU ping-pong and only 3.5% perplexity drop.

