HarvestBench Puts a Price on AI Compassion

π‘See which LLMs pay to avoid harmβand how strongly a simple moral briefing changes their actions.
β‘ 30-Second TL;DR
What Changed
The benchmark presents agents with a priced choice: drive over an animal for free or spend fuel to swerve.
Why It Matters
HarvestBench suggests that model capability alone is a poor predictor of safe agent behavior, especially when avoiding harm carries an explicit operational cost. Its results support evaluating agents through logged actions, incentives, and side effects rather than relying only on verbal safety judgments.
What To Do Next
Run your agent policy through HarvestBench or a similar priced side-effect simulation, and compare logged harm rates with and without explicit safety briefings.
Key Points
- β’The benchmark presents agents with a priced choice: drive over an animal for free or spend fuel to swerve.
- β’Across 7,201 decisions, animal kill rates ranged from 0.4% to 98.8% and were not correlated with model capability.
- β’Four of six tested models responded significantly to price, with measured price elasticities from 0.09 to 1.69.
- β’All nine models killed wild animals more often than farmed animals, while morality briefings reduced kill rates below 6% in five of six reasoning models.
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.