
Predicting Agent Coding Task Performance
Introduces framework using augmented Item Response Theory (IRT) to predict success/failure on individual agentic coding tasks. Decomposes agent ability into LLM and scaffold components for cross-leaderboard aggregation. Enables predictions for unseen benchmarks and agent combinations, aiding benchmark calibration.






