SourceStalecollected in 2h

LoCoMo Audit: 6.4% Key Errors, Judge Passes 63% Wrongs

PostLinkedIn
🤖Read original on Reddit r/MachineLearning
#benchmark-audit#long-context#llm-evaluationlocomolocomogpt-4o-minilongmemeval-s

💡LoCoMo flawed: 6.4% key errors, judge OKs 63% wrongs—rethink memory benchmarks now

⚡ 30-Second TL;DR

What Changed

99 errors in answer key: hallucinations, temporal reasoning, speaker attribution

Why It Matters

Exposes flaws in popular long-context memory benchmarks, urging caution in leaderboard comparisons and pushing for better alternatives.

What To Do Next

Download locomo-audit repo and validate your long-memory model scores against documented fixes.

Who should care:Researchers & Academics

Key Points

  • 99 errors in answer key: hallucinations, temporal reasoning, speaker attribution
  • gpt-4o-mini judge passes 62.81% intentionally wrong but topically adjacent answers
  • No standardized evaluation pipeline leads to irreproducible scores
  • Full audit repo with documented errors and reproducible scripts available
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.