ik_llama.cpp 26x Faster Qwen 3.5 Prompts
ik_llama.cpp fork delivers 26x faster prompt evaluation (43 to 1,122 tok/s) and 3.5x generation speed on Qwen 3.5 27B Q4_K_M using RTX PRO 4000. Fused GDN kernels reduce graph splits from 34 to 2 for full GPU utilization. Pre-built Windows binaries available as drop-in replacement.





