📄較早收集於 5h

BTF-2 基準測試 AI 預測代理推理

BTF-2 基準測試 AI 預測代理推理
PostLinkedIn
📄閱讀原文: ArXiv AI

💡新基準揭露 AI 預測者的策略缺陷—代理建構者必讀 (24字)

⚡ 30-Second TL;DR

有什麼變化

1,417 個過去預測問題與凍結 15M 文件語料庫,用於離線代理測試

為什麼重要

超越排行榜提供 AI 預測更深層分析,揭示特定策略弱點。引導代理在政治與商業等高風險領域改進。加速前沿模型更可靠推理的開發。

下一步行動

從 arXiv:2604.26106 下載 BTF-2,並在過去預測問題上基準測試你的預測代理。

誰應關注:Researchers & Academics

關鍵要點

  • 1,417 個過去預測問題與凍結 15M 文件語料庫,用於離線代理測試
  • 偵測 0.004 Brier 分數差異;區分研究與判斷優勢
  • 透過預 mortem 與黑天鵝分析建構 0.011 Brier 更優預測器
  • 前沿代理在領導者動機、計劃執行與機構流程上失敗

🧠 深度解析

AI-generated analysis for this event.

🔑 增強重點摘要

  • BTF-2 utilizes a 'frozen' corpus approach to solve the data contamination problem prevalent in LLM evaluation, ensuring that agents cannot access information published after the event date.
  • The benchmark introduces a novel 'Counterfactual Sensitivity' metric that measures how agents adjust their probability estimates when provided with specific, injected adversarial scenarios.
  • The research team behind BTF-2 identified that frontier models exhibit a 'confirmation bias' in forecasting, where they disproportionately weight evidence supporting their initial hypothesis over contradictory data.
📊 競品分析▸ Show
FeatureBTF-2Metaculus (Platform)Manifold Markets
Primary UseOffline Agent BenchmarkingHuman/Hybrid ForecastingPrediction Markets
Data SourceFrozen 15M-doc CorpusReal-time Web/Human InputReal-time Market Data
EvaluationAutomated Brier ScoreCrowd ConsensusMarket Price
PricingOpen Research/FreeFreemiumTransaction-based

🛠️ 技術深入

  • Corpus Architecture: Uses a static, time-stamped snapshot of 15 million documents (news, policy papers, financial reports) indexed via a vector database to simulate a 'closed-world' information environment.
  • Evaluation Engine: Implements a multi-stage pipeline that separates 'Information Retrieval' (RAG performance) from 'Probabilistic Reasoning' (Brier score calculation) to isolate failure points.
  • Agent Prompting: Employs a chain-of-thought (CoT) framework specifically tuned for 'Pre-Mortem' analysis, forcing agents to generate three distinct failure scenarios before outputting a final probability distribution.
  • Sensitivity Analysis: Uses a perturbation-based testing method where key variables in the prompt are modified to check for non-linear shifts in the agent's confidence intervals.

🔮 前景展望AI analysis grounded in cited sources

Standardized forecasting benchmarks will become a mandatory component of AI safety evaluations.
Regulators are increasingly requiring quantitative evidence of an AI's ability to model complex, long-term strategic outcomes as a condition for deployment.
Future LLM architectures will incorporate dedicated 'probabilistic reasoning' modules.
The failure of current frontier models to handle institutional incentives suggests that general-purpose transformers require specialized layers for strategic game theory.

時間線

2024-09
Initial release of BTF-1 focusing on basic geopolitical forecasting.
2025-06
Expansion of the frozen corpus to include specialized financial and scientific datasets.
2026-04
Official publication of BTF-2 with the 1,417-question dataset.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI