DFlash Boosts Qwen 35B 33% on 8GB GPU
User achieved DFlash speculative decoding on Qwen3.5-35B-A3B using llama.cpp on RTX 2080 SUPER 8GB VRAM via MoE CPU offload. Speed improved from 26.8 to 35.6 tok/s, a 33% gain. Optimal draft-max was 6 with 99% acceptance.