來源Reddit r/LocalLLaMA•較早收集於 4h
Claude 無法通關 Elden Ring:AGI 還遠
#agi-debate#llm-benchmarks#gaming-testclaudeclaudejensen-huangmarc-andreessenelden-ring
💡用 Claude 遊戲失敗駁斥 AGI 炒作—基準現實主義者必讀(38字)
⚡ 30 秒速覽
有什麼變化
批判 Jensen Huang 與 Marc Andreessen 的 AGI 宣稱
為什麼重要
引發 AGI 基準辯論,敦促從業者測試 LLM 於新型任務。
下一步行動
測試你的 LLM 於零樣本遊戲任務,如 Elden Ring 導航。
誰應關注:Researchers & Academics
關鍵要點
- •批判 Jensen Huang 與 Marc Andreessen 的 AGI 宣稱
- •Claude Opus 無法離開 Elden Ring 角色創建
- •強調超出分佈推理需求
- •發布於 r/LocalLLaMA 社群
🧠 深度解析
本篇為 AI 生成分析,非原文內容。
🔑 增強重點摘要
- •The failure of LLMs in complex, real-time environments like Elden Ring highlights the 'embodiment gap,' where models struggle with high-latency, non-deterministic visual feedback loops compared to static text-based reasoning.
- •Industry researchers distinguish between 'System 1' (fast, intuitive) and 'System 2' (slow, deliberative) reasoning; current architectures like Claude's struggle to maintain long-horizon planning in dynamic game environments without explicit neuro-symbolic integration.
- •The Reddit discourse reflects a broader shift in the AI community toward 'benchmarking by frustration,' where users test models against complex, multi-modal tasks to expose the limitations of current scaling laws.
📊 競品分析▸ Show
| Feature | Claude 3.5 Opus | GPT-4o | Gemini 1.5 Pro |
|---|---|---|---|
| Reasoning Architecture | Transformer-based (CoT) | Multimodal Transformer | Mixture-of-Experts |
| Context Window | 200k tokens | 128k tokens | 2M tokens |
| Game/Real-time Task Capability | Low (Text-heavy) | Low (Vision-limited) | Moderate (Long-context) |
| Pricing | $15/million input tokens | $5/million input tokens | $3.50/million input tokens |
🛠️ 技術深入
- •Current LLM architectures lack a persistent 'world model' state, preventing them from maintaining spatial awareness in 3D environments like Elden Ring.
- •The failure to exit the room is attributed to the lack of a closed-loop feedback mechanism; the model receives a frame, but cannot predict the consequences of its actions (e.g., 'press W') on the game state.
- •Claude Opus utilizes a standard Transformer decoder architecture optimized for text and code, which lacks the temporal memory required for continuous, real-time decision-making in gaming engines.
🔮 前景展望基於引用來源的 AI 分析
LLM-based agents will require dedicated 'World Model' layers to succeed in interactive gaming.
Without internal representations of physics and spatial constraints, models cannot perform the multi-step planning required for complex game navigation.
AGI definitions will shift from 'passing benchmarks' to 'demonstrating autonomous task completion in open-world environments'.
The failure in Elden Ring serves as a public litmus test that exposes the gap between high-scoring benchmarks and real-world utility.
⏳ 時間線
2024-03
Anthropic releases Claude 3 Opus, setting new industry benchmarks for reasoning.
2024-10
Anthropic releases Claude 3.5 Sonnet, focusing on improved agentic capabilities.
2025-06
Anthropic introduces 'Computer Use' capabilities, allowing models to interact with desktop interfaces.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Reddit r/LocalLLaMA ↗
每週電子報
每週一封,可隨時退訂。