📄ArXiv AI•較早收集於 17h
LLM-WikiRace 揭示 LLM 規劃限制
#benchmark#planning#reasoning#knowledge-graphsllm-wikirace
💡New benchmark shows even Gemini-3/GPT-5 fail 77% on hard LLM planning tasks
⚡ 30-Second TL;DR
有什麼變化
引入維基百科超連結導航任務測試 LLM 規劃
為什麼重要
此基準暴露前沿 LLM 推理的關鍵缺口,推動朝向更好長期規劃的開發。它提供簡單、可重現的測試環境,用於改善真實世界知識導航的代理能力。
下一步行動
Visit https://llmwikirace.github.io to download code and benchmark your LLM on hard levels.
誰應關注:Researchers & Academics
關鍵要點
- •引入維基百科超連結導航任務測試 LLM 規劃
- •Gemini-3 在困難層級達 23% 領先,簡單層級超人類
- •世界知識必要但長期規劃為主導因素
- •頂尖模型錯誤後進入迴圈而非重新規劃
- •開放程式碼與排行榜於 llmwikirace.github.io
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 2 個來源。
🔑 增強重點摘要
- •LLM-WikiRace benchmarks LLMs on navigating Wikipedia hyperlinks from source to target pages, stratifying tasks into easy, medium, and hard splits based on shortest-path length[1][2].
- •Frontier models like Gemini-3, GPT-5, and Claude Opus 4.5 achieve over 90% success on easy levels (superhuman performance) but drop to 23% for Gemini-3 on hard levels[1][2].
- •World knowledge is essential up to a threshold, after which long-horizon planning and reasoning dominate performance[1][2].
- •Top models fail to replan after errors, often entering loops instead of recovering[1][2].
- •Metrics include success rate, suboptimal steps (excess beyond shortest path), and average cost (tokens and monetary for closed models); step limit is 30, providing 3x budget over optimal paths up to 8 steps[1].
📊 競品分析▸ Show
| Feature | LLM-WikiRace | Other Benchmarks |
|---|---|---|
| Task | Wikipedia hyperlink navigation for planning/reasoning | Varies (e.g., multi-modal in others) |
| Difficulty Splits | Easy (>90% success), Medium (50-70%), Hard (<25%) | Not specified |
| Metrics | Success rate, suboptimal steps, avg cost | Varies |
| Frontier Model Hard Success | Gemini-3: 23% | N/A (isolates textual planning) |
🛠️ 技術深入
- •Task requires step-by-step hyperlink navigation with 30-step limit (plateaus beyond; longest optimal path is 8 steps)[1].
- •Algorithm detailed in Appendix B; evaluates full episodes including failures[1].
- •Fine-tuning improves easy split (22.5% to 67.5% after 300 steps), modest on medium (1.3% to 4.6%), none on hard (0%)[1].
- •Open code and leaderboard at llmwikirace.github.io[1][2].
- •Subjects: Artificial Intelligence (cs.AI), Machine Learning (cs.LG)[2].
🔮 前景展望AI analysis grounded in cited sources
Highlights persistent gaps in long-horizon planning for frontier LLMs, emphasizing need for better replanning and recovery mechanisms; serves as open benchmark to drive progress in reasoning systems beyond world knowledge reliance.
⏳ 時間線
2026-02
LLM-WikiRace paper released on arXiv (2602.16902v1), introducing benchmark with results on Gemini-3, GPT-5, Claude Opus 4.5
📎 來源 (2)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI ↗
每週 AI 簡報
每週一封,可隨時退訂。