LoCoMo Audit: 6.4% Key Errors, Judge Passes 63% Wrongs
💡LoCoMo flawed: 6.4% key errors, judge OKs 63% wrongs—rethink memory benchmarks now
⚡ 30-Second TL;DR
What Changed
99 errors in answer key: hallucinations, temporal reasoning, speaker attribution
Why It Matters
Exposes flaws in popular long-context memory benchmarks, urging caution in leaderboard comparisons and pushing for better alternatives.
What To Do Next
Download locomo-audit repo and validate your long-memory model scores against documented fixes.
Key Points
- •99 errors in answer key: hallucinations, temporal reasoning, speaker attribution
- •gpt-4o-mini judge passes 62.81% intentionally wrong but topically adjacent answers
- •No standardized evaluation pipeline leads to irreproducible scores
- •Full audit repo with documented errors and reproducible scripts available
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.