來源較早收集於 4h

Claude 無法通關 Elden Ring:AGI 還遠

PostLinkedIn
🦙閱讀原文: Reddit r/LocalLLaMA
#agi-debate#llm-benchmarks#gaming-testclaudeclaudejensen-huangmarc-andreessenelden-ring

💡用 Claude 遊戲失敗駁斥 AGI 炒作—基準現實主義者必讀(38字)

⚡ 30 秒速覽

有什麼變化

批判 Jensen Huang 與 Marc Andreessen 的 AGI 宣稱

為什麼重要

引發 AGI 基準辯論,敦促從業者測試 LLM 於新型任務。

下一步行動

測試你的 LLM 於零樣本遊戲任務,如 Elden Ring 導航。

誰應關注:Researchers & Academics

關鍵要點

  • 批判 Jensen Huang 與 Marc Andreessen 的 AGI 宣稱
  • Claude Opus 無法離開 Elden Ring 角色創建
  • 強調超出分佈推理需求
  • 發布於 r/LocalLLaMA 社群

🧠 深度解析

本篇為 AI 生成分析,非原文內容。

🔑 增強重點摘要

  • The failure of LLMs in complex, real-time environments like Elden Ring highlights the 'embodiment gap,' where models struggle with high-latency, non-deterministic visual feedback loops compared to static text-based reasoning.
  • Industry researchers distinguish between 'System 1' (fast, intuitive) and 'System 2' (slow, deliberative) reasoning; current architectures like Claude's struggle to maintain long-horizon planning in dynamic game environments without explicit neuro-symbolic integration.
  • The Reddit discourse reflects a broader shift in the AI community toward 'benchmarking by frustration,' where users test models against complex, multi-modal tasks to expose the limitations of current scaling laws.
📊 競品分析▸ Show
FeatureClaude 3.5 OpusGPT-4oGemini 1.5 Pro
Reasoning ArchitectureTransformer-based (CoT)Multimodal TransformerMixture-of-Experts
Context Window200k tokens128k tokens2M tokens
Game/Real-time Task CapabilityLow (Text-heavy)Low (Vision-limited)Moderate (Long-context)
Pricing$15/million input tokens$5/million input tokens$3.50/million input tokens

🛠️ 技術深入

  • Current LLM architectures lack a persistent 'world model' state, preventing them from maintaining spatial awareness in 3D environments like Elden Ring.
  • The failure to exit the room is attributed to the lack of a closed-loop feedback mechanism; the model receives a frame, but cannot predict the consequences of its actions (e.g., 'press W') on the game state.
  • Claude Opus utilizes a standard Transformer decoder architecture optimized for text and code, which lacks the temporal memory required for continuous, real-time decision-making in gaming engines.

🔮 前景展望基於引用來源的 AI 分析

LLM-based agents will require dedicated 'World Model' layers to succeed in interactive gaming.
Without internal representations of physics and spatial constraints, models cannot perform the multi-step planning required for complex game navigation.
AGI definitions will shift from 'passing benchmarks' to 'demonstrating autonomous task completion in open-world environments'.
The failure in Elden Ring serves as a public litmus test that exposes the gap between high-scoring benchmarks and real-world utility.

時間線

2024-03
Anthropic releases Claude 3 Opus, setting new industry benchmarks for reasoning.
2024-10
Anthropic releases Claude 3.5 Sonnet, focusing on improved agentic capabilities.
2025-06
Anthropic introduces 'Computer Use' capabilities, allowing models to interact with desktop interfaces.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Reddit r/LocalLLaMA

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週電子報

每週一封,可隨時退訂。