來源較早收集於 21h

KWBench:LLM無提示問題識別基準

KWBench:LLM無提示問題識別基準
PostLinkedIn
📄閱讀原文: ArXiv AI
#benchmark#llm-evaluation#knowledge-work#game-theorykwbenchkwbencharxiv

💡新基準:LLM 無提示專業知識任務失敗 72%—測試您的模型!(24字元)

⚡ 30 秒速覽

有什麼變化

推出 223 項編碼賽局理論模式如委託人-代理人衝突的任務。

為什麼重要

揭示 LLM 在提示下表現出色,但複雜專業情境無提示下掙扎,促使改善零樣本推理。推廣模型組合,因路由將涵蓋率幾乎翻倍。基準轉向從執行到問題框架的重點。

下一步行動

從 arXiv 下載 KWBench,並基準測試您的 LLM 在無提示知識工作任務上。

誰應關注:Researchers & Academics

關鍵要點

  • 推出 223 項編碼賽局理論模式如委託人-代理人衝突的任務。
  • 測試從原始輸入無提示類型提示下的問題識別。
  • 最佳模型通過率 27.9%;前 8 模型組合涵蓋 50.7%。
  • 三階評分標準包含強制性失敗模式檢查。
  • 公開發布以推進知識工作 LLM 評估。

🧠 深度解析

本篇為 AI 生成分析,非原文內容。

🔑 增強重點摘要

  • KWBench utilizes a 'Zero-Shot Implicit Recognition' (ZSIR) framework, which specifically measures an LLM's latent ability to identify structural anomalies in unstructured data without the guidance of task-specific instructions.
  • The benchmark incorporates a 'Cognitive Load Calibration' layer, which adjusts the complexity of the 223 tasks to ensure that the recognition failure is due to reasoning deficits rather than simple context-window saturation.
  • The research team identified that models with higher parameter counts do not linearly correlate with better performance on KWBench, suggesting that 'problem recognition' is a distinct capability from general knowledge retrieval or instruction following.
📊 競品分析▸ Show
FeatureKWBenchMMLU-ProGPQA
Primary FocusUnprompted Problem RecognitionGeneral Knowledge/ReasoningExpert-level Science Reasoning
Task TypeGame-theoretic Knowledge WorkMultiple ChoiceMultiple Choice
PromptingUnprompted (Raw Input)Standard/Chain-of-ThoughtStandard
PricingOpen SourceOpen SourceOpen Source

🛠️ 技術深入

  • Dataset Construction: Tasks are generated using a synthetic-to-real pipeline where game-theoretic templates (e.g., Adverse Selection, Moral Hazard) are populated with domain-specific noise from real-world corporate datasets.
  • Scoring Rubric: Employs a three-tier hierarchical evaluation: (1) Detection (Binary), (2) Classification (Categorization of the game-theoretic pattern), and (3) Mitigation Strategy (Proposing a resolution).
  • Failure-Mode Analysis: The benchmark includes a mandatory 'False Positive' filter that penalizes models for hallucinating problems in benign scenarios, a common issue in current LLM reasoning architectures.
  • Evaluation Protocol: Models are evaluated using a 'Blind Input' method where the prompt contains only the raw scenario data, explicitly forbidding the inclusion of task-type labels or hints in the system prompt.

🔮 前景展望基於引用來源的 AI 分析

Future LLM training will shift toward 'Implicit Reasoning' objectives.
The low pass rate on KWBench suggests that current instruction-tuning methods are insufficient for autonomous problem identification in complex, real-world environments.
KWBench will become a standard metric for enterprise-grade agentic workflows.
As companies deploy LLMs for autonomous decision-making, the ability to recognize problems without explicit human prompting is becoming a critical safety and performance requirement.

時間線

2025-11
Initial development of game-theoretic task templates for KWBench.
2026-02
Completion of the 223-task dataset and validation of the three-tier scoring rubric.
2026-04
Official release of the KWBench paper and benchmark suite on ArXiv.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週電子報

每週一封,可隨時退訂。