Google: Longer CoT Hurts Accuracy (-0.54)
Google's paper reveals longer chain-of-thought reasoning correlates negatively with accuracy (-0.54) across models like GPT-OSS and Qwen3. They introduce DTR to measure deep thinking fraction and Think@n strategy for efficient sampling, cutting compute by 50% with better results.
Reddit r/LocalLLaMA · 203d ago


















