Qwen 35B-A3B: 26 t/s on 8GB Laptop at 100K Context
Qwen3.5-35B-A3B-UD-Q4_K_XL on RTX 4060 8GB gaming laptop with llama.cpp achieves 26 t/s generation at 100K context and 331 t/s prompt processing. Performance scales well from 5K to 100K contexts. User eyes RX 7900 XTX upgrade for larger models.





