🧠較早收集於 61m

DeepMind 駁斥「越多智能體越好」

DeepMind 駁斥「越多智能體越好」
PostLinkedIn
🧠閱讀原文: 机器之心
#multi-agent#scaling-laws#agent-benchmarksdeepmind-agent-scalingdeepmindgptgeminiclaude

💡DeepMind: More agents often worsen perf—first scaling laws for agent systems from 180 evals

⚡ 30-Second TL;DR

有什麼變化

測試 180 種配置:單體 vs 多智能體(獨立、集中式等)

為什麼重要

挑戰多智能體炒作,引導助手及規劃器等真實 AI 應用更好設計。

下一步行動

Read arXiv 2512.08296 and test centralized agents on Finance-Agent benchmark for your workflows.

誰應關注:Researchers & Academics

關鍵要點

  • 測試 180 種配置:單體 vs 多智能體(獨立、集中式等)
  • 智能體增多遇瓶頸,甚至降性能
  • 任務屬性決定最佳智能體數/架構
  • 基準:Finance-Agent、BrowseComp-Plus、PlanCraft、Workbench
  • 論文:Towards a Science of Scaling Agent Systems

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 4 個來源。

🔑 增強重點摘要

  • The predictive model achieves cross-validated R²=0.524 using coordination metrics like efficiency, overhead, error amplification, and redundancy to forecast performance on unseen tasks.[1][2]
  • Independent agents amplify errors 17.2x compared to centralized coordination's 4.4x, highlighting topology-dependent error propagation.[1]
  • Out-of-sample validation on frontier models like GPT-5.2, Gemini-3.0 Pro, and Flash confirms four of five scaling principles with MAE=0.071-0.077.[1][2]

🛠️ 技術深入

  • Five canonical architectures: Single-Agent, Independent Multi-Agent, Centralized (80.8% gain on parallel tasks), Decentralized (+9.2% on web navigation), Hybrid.[1][2][4]
  • Controlled setup standardizes tools, prompts, and token budgets across 180 configs using three LLM families (GPT, Gemini, Claude) to isolate architecture effects.[1][3]
  • Three key effects: tool-coordination trade-off (tool-heavy tasks suffer multi-agent overhead), capability saturation (diminishing returns above ~45% single-agent baseline), topology-dependent error amplification.[1]

🔮 前景展望AI analysis grounded in cited sources

Agent design will shift from heuristic 'more agents' to task-property-based prediction models achieving 87% optimal strategy accuracy.
The framework's predictive model uses measurable task properties like tool count and decomposability to select architectures for unseen configurations.[1][2][4]
Multi-agent systems will integrate with advancing frontier models like Gemini-3.0 without replacing single-agent baselines on sequential tasks.
Validation on GPT-5.2 and Gemini-3.0 confirms scaling principles generalize, but multi-agent degrades sequential reasoning by 39-70%.[1][2][4]

時間線

2025-12
Paper 'Towards a Science of Scaling Agent Systems' submitted to arXiv (v1: Dec 9, v2: Dec 17)
2026-01
Google Research blog post published detailing 180-config evaluation and predictive model
2026-02
机器之心 article summarizes DeepMind paper on agent scaling limits

📎 來源 (4)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. arXiv — 2512
  2. sciety.org — V1
  3. arXiv — 2512
  4. research.google — Towards a Science of Scaling Agent Systems When and Why Agent Systems Work
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 机器之心

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。