
Llama.cpp Adds Real Reasoning Budget Control
Llama.cpp introduces sampler-based reasoning budget to limit thinking tokens and force termination. Includes --reasoning-budget-message flag to smooth transitions, restoring performance on Qwen3 9B HumanEval from 78% to 89%. Experimentation encouraged for optimal settings.




