
AI Debate Training Curbs Reward Hacking
Google DeepMind researchers found that training AI agents through debate can reduce reward hacking when an LLM judge evaluates their work. On mathematics tasks, debate achieved higher sustained ground-truth accuracy than directly optimizing for LLM judge rewards.






