BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces

๐กStop relying on synthetic users; use this real-world benchmark to test how well your AI models understand human behavior
โก 30-Second TL;DR
What Changed
Features 141,445 belief prediction instances and 1,485,972 trade prediction instances.
Why It Matters
This benchmark provides a more rigorous standard for personalized AI, helping researchers identify failure modes in how models interpret human behavior. It shifts the focus from simulated benchmarks to verifiable, real-world decision-making patterns.
What To Do Next
Download the BehaviorBench dataset from the ArXiv repository to benchmark your personalized recommendation or agentic decision-making models against real-world wallet traces.
Key Points
- โขFeatures 141,445 belief prediction instances and 1,485,972 trade prediction instances.
- โขUses real-world wallet-level decision histories instead of synthetic or model-generated data.
- โขEvaluates four history interfaces: no personalization, recent history, generated profiles, and retrieved evidence.
- โขDemonstrates that personalization impacts belief prediction more significantly than transaction prediction.
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI โ