📄較早收集於 17h

LLM-WikiRace 揭示 LLM 規劃限制

LLM-WikiRace 揭示 LLM 規劃限制
PostLinkedIn
📄閱讀原文: ArXiv AI
#benchmark#planning#reasoning#knowledge-graphsllm-wikirace

💡New benchmark shows even Gemini-3/GPT-5 fail 77% on hard LLM planning tasks

⚡ 30-Second TL;DR

有什麼變化

引入維基百科超連結導航任務測試 LLM 規劃

為什麼重要

此基準暴露前沿 LLM 推理的關鍵缺口,推動朝向更好長期規劃的開發。它提供簡單、可重現的測試環境,用於改善真實世界知識導航的代理能力。

下一步行動

Visit https://llmwikirace.github.io to download code and benchmark your LLM on hard levels.

誰應關注:Researchers & Academics

關鍵要點

  • 引入維基百科超連結導航任務測試 LLM 規劃
  • Gemini-3 在困難層級達 23% 領先,簡單層級超人類
  • 世界知識必要但長期規劃為主導因素
  • 頂尖模型錯誤後進入迴圈而非重新規劃
  • 開放程式碼與排行榜於 llmwikirace.github.io

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 2 個來源。

🔑 增強重點摘要

  • LLM-WikiRace benchmarks LLMs on navigating Wikipedia hyperlinks from source to target pages, stratifying tasks into easy, medium, and hard splits based on shortest-path length[1][2].
  • Frontier models like Gemini-3, GPT-5, and Claude Opus 4.5 achieve over 90% success on easy levels (superhuman performance) but drop to 23% for Gemini-3 on hard levels[1][2].
  • World knowledge is essential up to a threshold, after which long-horizon planning and reasoning dominate performance[1][2].
  • Top models fail to replan after errors, often entering loops instead of recovering[1][2].
  • Metrics include success rate, suboptimal steps (excess beyond shortest path), and average cost (tokens and monetary for closed models); step limit is 30, providing 3x budget over optimal paths up to 8 steps[1].
📊 競品分析▸ Show
FeatureLLM-WikiRaceOther Benchmarks
TaskWikipedia hyperlink navigation for planning/reasoningVaries (e.g., multi-modal in others)
Difficulty SplitsEasy (>90% success), Medium (50-70%), Hard (<25%)Not specified
MetricsSuccess rate, suboptimal steps, avg costVaries
Frontier Model Hard SuccessGemini-3: 23%N/A (isolates textual planning)

🛠️ 技術深入

  • Task requires step-by-step hyperlink navigation with 30-step limit (plateaus beyond; longest optimal path is 8 steps)[1].
  • Algorithm detailed in Appendix B; evaluates full episodes including failures[1].
  • Fine-tuning improves easy split (22.5% to 67.5% after 300 steps), modest on medium (1.3% to 4.6%), none on hard (0%)[1].
  • Open code and leaderboard at llmwikirace.github.io[1][2].
  • Subjects: Artificial Intelligence (cs.AI), Machine Learning (cs.LG)[2].

🔮 前景展望AI analysis grounded in cited sources

Highlights persistent gaps in long-horizon planning for frontier LLMs, emphasizing need for better replanning and recovery mechanisms; serves as open benchmark to drive progress in reasoning systems beyond world knowledge reliance.

時間線

2026-02
LLM-WikiRace paper released on arXiv (2602.16902v1), introducing benchmark with results on Gemini-3, GPT-5, Claude Opus 4.5

📎 來源 (2)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. arXiv — 2602
  2. arXiv — 2602
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。