來源ArXiv AI•較早收集於 3h
XpertBench:專家級 LLM 基準測試

#benchmark#rubrics#expert-tasks#llm-evaluationxpertbenchxpertbenchshotjudge
💡新基準揭露 LLM 專家差距:頂尖模型僅 55%—評估必讀!
⚡ 30 秒速覽
有什麼變化
1,346 個任務來自 80 類別,由 1,000 多位領域專家策劃
為什麼重要
XpertBench 提升 LLM 評估標準,揭示當前模型在專家認知上的限制,並推動專門化 AI 開發。它提供可擴展且人類對齊的工具,用以追蹤邁向專業級助手的進展。
下一步行動
從 arXiv:2604.02368 下載 XpertBench 任務並基準測試您的 LLM。
誰應關注:Researchers & Academics
關鍵要點
- •1,346 個任務來自 80 類別,由 1,000 多位領域專家策劃
- •評分表包含 15-40 個加權檢查點,用於專業評估
- •ShotJudge 使用少樣本範例校準的 LLM 評判者,避免偏差
- •頂尖 LLM 平均 55% 分數,最高 66%,領域表現分歧
🧠 深度解析
本篇為 AI 生成分析,非原文內容。
🔑 增強重點摘要
- •XpertBench utilizes a dynamic 'Difficulty-Weighted Scoring' (DWS) mechanism that adjusts rubric checkpoints based on the historical failure rates of previous SOTA models, preventing score saturation.
- •The benchmark includes a 'Cross-Domain Consistency' metric, which specifically measures if models maintain reasoning integrity when presented with the same logic problem across different professional contexts (e.g., legal vs. medical).
- •The ShotJudge framework incorporates a 'Self-Correction Loop' where the judge model is required to generate a critique of its own initial assessment before finalizing the score, significantly reducing the 'length bias' common in LLM-as-a-judge systems.
📊 競品分析▸ Show
| Feature | XpertBench | MMLU-Pro | GPQA | HumanEval |
|---|---|---|---|---|
| Focus | Professional Expert Tasks | Advanced Reasoning | PhD-level Science | Coding |
| Judging | ShotJudge (Few-shot) | GPT-4o / Rule-based | Expert Human | Unit Tests |
| Scale | 1,346 Tasks | 12,000+ Questions | 448 Questions | 164 Problems |
| Pricing | Open Source (Benchmark) | Open Source | Open Source | Open Source |
🛠️ 技術深入
- •Rubric Architecture: Each task employs a hierarchical rubric structure where 15-40 checkpoints are categorized into 'Core Accuracy,' 'Professional Nuance,' and 'Regulatory Compliance'.
- •ShotJudge Implementation: Utilizes a 5-shot prompt template containing high-variance examples (correct, incorrect, and partially correct) to calibrate the judge model's latent scoring distribution.
- •Bias Mitigation: Employs a 'Position-Balanced' evaluation strategy where the judge evaluates the model output in both original and reversed order to mitigate positional bias.
- •Data Integrity: The dataset is hosted on a version-controlled repository with a 'Contamination-Detection' layer that cross-references task prompts against common pre-training corpora (e.g., Common Crawl, Pile) to ensure zero-shot validity.
🔮 前景展望基於引用來源的 AI 分析
XpertBench will become the primary standard for enterprise-grade LLM procurement by 2027.
The focus on professional domain-specific rubrics addresses the current lack of industry-standard metrics for high-stakes business deployment.
Model developers will shift training focus toward 'Expert-Gap' reduction.
The 66% peak success rate highlights a significant performance ceiling that will force architectural changes in reasoning-heavy models.
⏳ 時間線
2025-09
Initial pilot phase of XpertBench involving 200 tasks and 150 domain experts.
2026-01
Expansion of the expert panel to 1,000+ contributors and finalization of the 80-domain taxonomy.
2026-04
Official release of the XpertBench dataset and ShotJudge framework on ArXiv.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI ↗
每週電子報
每週一封,可隨時退訂。