
AgentSelect Benchmark for Agent Recommendation
AgentSelect introduces a benchmark for recommending LLM agent configurations based on narrative queries, addressing the lack of query-conditioned supervision. It aggregates 111,179 queries, 107,721 agents, and 251,103 interactions from 40+ sources into unified data. Analyses highlight the shift to long-tail supervision and the need for content-aware capability matching.

