📄較早收集於 12h

代理狀態評估實現 LLM 代理基準可擴展

代理狀態評估實現 LLM 代理基準可擴展
PostLinkedIn
📄閱讀原文: ArXiv AI
#multi-turn-agents#tool-calling#llm-judges#simulation-benchmarkproxy-state-based-evaluation

💡Scalable agent eval framework: 90%+ agreement, no costly DBs, beats tau-bench setups

⚡ 30-Second TL;DR

有什麼變化

LLM 狀態追蹤器從完整互動軌跡推斷結構化代理狀態

為什麼重要

此框架降低建置代理基準的門檻,加速生產 LLM 代理開發。它支援訓練用的 on-policy 資料及使用者角色敏感度分析,有益產業應用。

下一步行動

Test Proxy State-Based Evaluation on your multi-turn agent benchmarks using LLM trackers for state inference.

誰應關注:Researchers & Academics

關鍵要點

  • LLM 狀態追蹤器從完整互動軌跡推斷結構化代理狀態
  • 透過 LLM 評判驗證目標完成並偵測工具/使用者幻覺
  • 產生跨代理家族的穩定、模型區分排名
  • 精心情境設計實現近零模擬器幻覺率
  • 人 LLM 評判一致率超過 90% 確保可靠評估

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 8 個來源。

🔑 增強重點摘要

  • Proxy State-Based Evaluation provides a scalable alternative to deterministic agentic benchmarks by using LLM-driven simulation, eliminating the engineering burden of maintaining fully deterministic backends[1]
  • The framework achieves consistent capability ordering across model families, with goal completion scaling predictably with model strength and inference-time reasoning effort[1]
  • Human-LLM judge agreement exceeds 90% with near-zero simulator hallucination rates, demonstrating reliable automated evaluation when scenarios are carefully specified[1]
  • The benchmark supports both on-policy and off-policy training data that transfers to unseen scenarios, enabling supervised learning improvements for open-weight reasoning agents[1]
  • Proxy state-based evaluation represents an emerging pattern in LLM agent benchmarking alongside complementary approaches like state-diff contracts for enterprise APIs and robustness testing under noisy conditions[2][3]
📊 競品分析▸ Show
ApproachEvaluation MethodHallucination RateJudge AgreementScalabilityUse Case
Proxy State-Based EvaluationLLM-driven simulation with state trackingNear-zero>90%High (no deterministic backend)Multi-turn tool-calling agents
State-Diff ContractsSandbox snapshots comparing initial/final statesN/AN/AHigh (isolated environments)Enterprise API tasks (224 tasks)
AgentNoiseBenchNoise injection with trajectory-aware evaluationN/AN/AHigh (automated pipeline)Robustness under adversarial conditions
AgentDAMWeb automation with contextual appropriateness framingN/A0.82-0.87 κModeratePrivacy leakage in multi-agent systems

🛠️ 技術深入

Scenario Schema: Each scenario specifies user goal, user/system facts, expected final state, and expected agent behavior, enabling structured evaluation without deterministic databases • LLM State Tracker Component: Infers structured proxy state from full interaction trace, preserving final state-based evaluation semantics • LLM Judge Verification: Verifies goal completion and detects tool/user hallucinations against scenario constraints with >90% agreement with human judges • Ablation Study Results: Confirms robustness of proxy state tracker and sensitivity to scenario completeness; user persona variability captured while maintaining low user-induced error • Model-Differentiating Rankings: Produces stable, interpretable metrics across model families and reasoning-effort settings (SFT, RFT training approaches) • Complementary Evaluation Patterns: Integrates with trajectory-aware protocols for multi-dimensional analysis and state-diff methodologies for outcome verification

🔮 前景展望AI analysis grounded in cited sources

Proxy State-Based Evaluation addresses a critical scalability bottleneck in LLM agent benchmarking by decoupling evaluation from deterministic backend maintenance. This framework enables rapid iteration on agent benchmarks without proportional infrastructure costs, likely accelerating the pace of agent capability assessment across industry. The >90% human-LLM judge agreement validates automated evaluation at scale, reducing evaluation costs while maintaining reliability. The transferability of training data to unseen scenarios suggests this approach could become a standard pattern for industrial LLM agent development, particularly as multi-agent systems and tool-calling complexity increase. Integration with complementary approaches like robustness testing under noise and privacy leakage detection indicates a maturing ecosystem of specialized benchmarking frameworks tailored to different agent deployment contexts.

時間線

2024-2025
Emergence of state-based evaluation methodologies for LLM agents, including state-diff contracts and proxy-guided approaches
2025-2026
Development of specialized benchmarking frameworks addressing robustness (AgentNoiseBench), privacy (AgentDAM), temporal reasoning (TemporalBench), and enterprise APIs
2026-02
Publication of Proxy State-Based Evaluation framework demonstrating >90% judge agreement and near-zero hallucination rates for multi-turn tool-calling agents

📎 來源 (8)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. arXiv — 2602
  2. arXiv — 2602
  3. arXiv — 2602
  4. arXiv — 2602
  5. arXiv — 2602
  6. arXiv — 2602
  7. arXiv — 2503
  8. philosophyofcomputing.substack.com — Coding Agents for Philosophers New
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。