GT-HarmBench: Game Theory AI Safety Benchmark
GT-HarmBench introduces 2,009 high-stakes multi-agent scenarios using game theory like Prisoner's Dilemma to benchmark AI safety risks. Frontier models select socially beneficial actions only 62% of the time, often leading to harm. The benchmark, code, and analysis are available on GitHub.






