Corecraft RL 環境訓練通用代理
💡RL env boosts agent pass@1 11% + OOD transfer up to 7.4%—key for enterprise AI
⚡ 30-Second TL;DR
有什麼變化
推出 Corecraft:客戶支援模擬,含 2,500 實體、14 類型、23 工具。
為什麼重要
如 Corecraft 般高品質 RL 環境證明對訓練超出訓練分佈泛化代理至關重要,滿足真實企業需求。此轉變強調從模型規模轉向環境設計,以實現可擴展代理能力。
下一步行動
Download Corecraft from EnterpriseGym and benchmark your RL agent on its tasks.
關鍵要點
- •推出 Corecraft:客戶支援模擬,含 2,500 實體、14 類型、23 工具。
- •前沿模型 (GPT-5.2、Claude Opus 4.6) 依專家標準僅解決 <30% 任務。
- •GRPO + 自適應裁剪訓練 GLM 4.6:25.37% → 36.76% 通過率,轉移至 OOD 基準。
- •任務導向建構、評分標準、企業工作流程促成泛化。
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 6 個來源。
🔑 增強重點摘要
- •Corecraft is the first environment in EnterpriseGym, Surge AI's suite of agentic RL environments, designed to train AI agents on realistic enterprise workflows with over 2,500 entities across 14 entity types and 23 unique tools[1][2]
- •Frontier models including GPT-5.2 and Claude Opus 4.6 achieve less than 30% task pass rate on Corecraft when all expert-authored rubric criteria must be satisfied, establishing a significant capability gap in current frontier models[2]
- •GLM 4.6 trained with Group Relative Policy Optimization (GRPO) and adaptive clipping improved from 25.37% to 36.76% task pass rate on held-out evaluation tasks after a single epoch of training[2]
- •Training gains on Corecraft transfer to out-of-distribution benchmarks with +4.5% improvement on BFCL Parallel, +7.4% on τ²-Bench Retail, and +6.8% on Toolathlon (Pass@1), demonstrating genuine generalization beyond the training distribution[2]
- •Three core design principles drive Corecraft's effectiveness: task-centric world building optimized for diverse and challenging tasks, expert-authored rubrics enabling reliable reward computation, and enterprise workflows reflecting realistic professional patterns[1][2]
🛠️ 技術深入
• Corecraft simulates a customer support agent at a fictional PC parts retailer (Corecraft Computers, Inc.), providing a stateful world where agents interact with databases, tools, and simulated customers[1] • Training methodology employs Group Relative Policy Optimization (GRPO) with adaptive clipping, representing an advancement in reinforcement learning training techniques for agentic systems[2] • Environment design prioritizes task quality and diversity over raw entity or tool counts, contrasting with approaches that maximize complexity without sufficient functional diversity[1] • Expert-authored rubrics provide structured evaluation criteria for task completion, enabling reliable reward signals during training[1][2] • The environment comprises 14 distinct entity types supporting multi-step, domain-specific work patterns typical of real enterprise customer support operations[1][2]
🔮 前景展望AI analysis grounded in cited sources
Corecraft establishes a new paradigm for training generalizable AI agents through high-fidelity, task-centric environments rather than synthetic or simplified training substrates. The demonstrated transfer to out-of-distribution benchmarks suggests that environment quality and realism are critical factors for developing agents capable of handling real-world enterprise workflows. This approach may influence how organizations develop and evaluate agentic AI systems, shifting focus from raw capability metrics to practical task completion in realistic scenarios. The success of GRPO training on Corecraft could accelerate adoption of similar high-fidelity simulation environments across other enterprise domains, potentially creating a new category of specialized RL environments for professional AI agent development.
📎 來源 (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI ↗
每週 AI 簡報
每週一封,可隨時退訂。
