SourceStalecollected in 7h

A New Map for Diagnosing Agent Failures

Read original on ArXiv AI
#agent-evaluation#failure-analysis#benchmarking#orchestration

Learn whether an agent failure needs model tuning, harness fixes, environment changes, or benchmark repair.

30-Second TL;DR

What Changed

The taxonomy assigns each of 41 failure modes to an interaction edge between two system components.

Why It Matters

The framework could reduce costly trial-and-error by helping teams distinguish model limitations from orchestration, tool, environment, or benchmark problems. It may also make agent evaluations more actionable and comparable across architectures.

What To Do Next

Annotate your agent logs by component-to-component interaction and fault side, then measure inter-rater agreement before prioritizing model fine-tuning.

Who should care:Researchers & Academics

Key Points

  • The taxonomy assigns each of 41 failure modes to an interaction edge between two system components.
  • A fault-side designation identifies which component should receive the repair or intervention.
  • The framework applies across coding assistants, long-horizon personal assistants, and multi-agent systems.
  • Across four frontier models, the strongest independent judge achieved Cohen's kappa of 0.76 against human labels.

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.