來源ArXiv AI•較早收集於 21h
KWBench:LLM無提示問題識別基準

#benchmark#llm-evaluation#knowledge-work#game-theorykwbenchkwbencharxiv
💡新基準:LLM 無提示專業知識任務失敗 72%—測試您的模型!(24字元)
⚡ 30 秒速覽
有什麼變化
推出 223 項編碼賽局理論模式如委託人-代理人衝突的任務。
為什麼重要
揭示 LLM 在提示下表現出色,但複雜專業情境無提示下掙扎,促使改善零樣本推理。推廣模型組合,因路由將涵蓋率幾乎翻倍。基準轉向從執行到問題框架的重點。
下一步行動
從 arXiv 下載 KWBench,並基準測試您的 LLM 在無提示知識工作任務上。
誰應關注:Researchers & Academics
關鍵要點
- •推出 223 項編碼賽局理論模式如委託人-代理人衝突的任務。
- •測試從原始輸入無提示類型提示下的問題識別。
- •最佳模型通過率 27.9%;前 8 模型組合涵蓋 50.7%。
- •三階評分標準包含強制性失敗模式檢查。
- •公開發布以推進知識工作 LLM 評估。
🧠 深度解析
本篇為 AI 生成分析,非原文內容。
🔑 增強重點摘要
- •KWBench utilizes a 'Zero-Shot Implicit Recognition' (ZSIR) framework, which specifically measures an LLM's latent ability to identify structural anomalies in unstructured data without the guidance of task-specific instructions.
- •The benchmark incorporates a 'Cognitive Load Calibration' layer, which adjusts the complexity of the 223 tasks to ensure that the recognition failure is due to reasoning deficits rather than simple context-window saturation.
- •The research team identified that models with higher parameter counts do not linearly correlate with better performance on KWBench, suggesting that 'problem recognition' is a distinct capability from general knowledge retrieval or instruction following.
📊 競品分析▸ Show
| Feature | KWBench | MMLU-Pro | GPQA |
|---|---|---|---|
| Primary Focus | Unprompted Problem Recognition | General Knowledge/Reasoning | Expert-level Science Reasoning |
| Task Type | Game-theoretic Knowledge Work | Multiple Choice | Multiple Choice |
| Prompting | Unprompted (Raw Input) | Standard/Chain-of-Thought | Standard |
| Pricing | Open Source | Open Source | Open Source |
🛠️ 技術深入
- •Dataset Construction: Tasks are generated using a synthetic-to-real pipeline where game-theoretic templates (e.g., Adverse Selection, Moral Hazard) are populated with domain-specific noise from real-world corporate datasets.
- •Scoring Rubric: Employs a three-tier hierarchical evaluation: (1) Detection (Binary), (2) Classification (Categorization of the game-theoretic pattern), and (3) Mitigation Strategy (Proposing a resolution).
- •Failure-Mode Analysis: The benchmark includes a mandatory 'False Positive' filter that penalizes models for hallucinating problems in benign scenarios, a common issue in current LLM reasoning architectures.
- •Evaluation Protocol: Models are evaluated using a 'Blind Input' method where the prompt contains only the raw scenario data, explicitly forbidding the inclusion of task-type labels or hints in the system prompt.
🔮 前景展望基於引用來源的 AI 分析
Future LLM training will shift toward 'Implicit Reasoning' objectives.
The low pass rate on KWBench suggests that current instruction-tuning methods are insufficient for autonomous problem identification in complex, real-world environments.
KWBench will become a standard metric for enterprise-grade agentic workflows.
As companies deploy LLMs for autonomous decision-making, the ability to recognize problems without explicit human prompting is becoming a critical safety and performance requirement.
⏳ 時間線
2025-11
Initial development of game-theoretic task templates for KWBench.
2026-02
Completion of the 223-task dataset and validation of the three-tier scoring rubric.
2026-04
Official release of the KWBench paper and benchmark suite on ArXiv.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI ↗
每週電子報
每週一封,可隨時退訂。