🧠机器之心•較早收集於 61m
DeepMind 駁斥「越多智能體越好」

#multi-agent#scaling-laws#agent-benchmarksdeepmind-agent-scalingdeepmindgptgeminiclaude
💡DeepMind: More agents often worsen perf—first scaling laws for agent systems from 180 evals
⚡ 30-Second TL;DR
有什麼變化
測試 180 種配置:單體 vs 多智能體(獨立、集中式等)
為什麼重要
挑戰多智能體炒作,引導助手及規劃器等真實 AI 應用更好設計。
下一步行動
Read arXiv 2512.08296 and test centralized agents on Finance-Agent benchmark for your workflows.
誰應關注:Researchers & Academics
關鍵要點
- •測試 180 種配置:單體 vs 多智能體(獨立、集中式等)
- •智能體增多遇瓶頸,甚至降性能
- •任務屬性決定最佳智能體數/架構
- •基準:Finance-Agent、BrowseComp-Plus、PlanCraft、Workbench
- •論文:Towards a Science of Scaling Agent Systems
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 4 個來源。
🔑 增強重點摘要
- •The predictive model achieves cross-validated R²=0.524 using coordination metrics like efficiency, overhead, error amplification, and redundancy to forecast performance on unseen tasks.[1][2]
- •Independent agents amplify errors 17.2x compared to centralized coordination's 4.4x, highlighting topology-dependent error propagation.[1]
- •Out-of-sample validation on frontier models like GPT-5.2, Gemini-3.0 Pro, and Flash confirms four of five scaling principles with MAE=0.071-0.077.[1][2]
🛠️ 技術深入
- •Five canonical architectures: Single-Agent, Independent Multi-Agent, Centralized (80.8% gain on parallel tasks), Decentralized (+9.2% on web navigation), Hybrid.[1][2][4]
- •Controlled setup standardizes tools, prompts, and token budgets across 180 configs using three LLM families (GPT, Gemini, Claude) to isolate architecture effects.[1][3]
- •Three key effects: tool-coordination trade-off (tool-heavy tasks suffer multi-agent overhead), capability saturation (diminishing returns above ~45% single-agent baseline), topology-dependent error amplification.[1]
🔮 前景展望AI analysis grounded in cited sources
Agent design will shift from heuristic 'more agents' to task-property-based prediction models achieving 87% optimal strategy accuracy.
⏳ 時間線
2025-12
Paper 'Towards a Science of Scaling Agent Systems' submitted to arXiv (v1: Dec 9, v2: Dec 17)
2026-01
Google Research blog post published detailing 180-config evaluation and predictive model
2026-02
机器之心 article summarizes DeepMind paper on agent scaling limits
📎 來源 (4)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 机器之心 ↗
每週 AI 簡報
每週一封,可隨時退訂。