A New Map for Diagnosing Agent Failures

Learn whether an agent failure needs model tuning, harness fixes, environment changes, or benchmark repair.
30-Second TL;DR
What Changed
The taxonomy assigns each of 41 failure modes to an interaction edge between two system components.
Why It Matters
The framework could reduce costly trial-and-error by helping teams distinguish model limitations from orchestration, tool, environment, or benchmark problems. It may also make agent evaluations more actionable and comparable across architectures.
What To Do Next
Annotate your agent logs by component-to-component interaction and fault side, then measure inter-rater agreement before prioritizing model fine-tuning.
Key Points
- •The taxonomy assigns each of 41 failure modes to an interaction edge between two system components.
- •A fault-side designation identifies which component should receive the repair or intervention.
- •The framework applies across coding assistants, long-horizon personal assistants, and multi-agent systems.
- •Across four frontier models, the strongest independent judge achieved Cohen's kappa of 0.76 against human labels.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.