📄ArXiv AI•較早收集於 5h
BTF-2 基準測試 AI 預測代理推理

💡新基準揭露 AI 預測者的策略缺陷—代理建構者必讀 (24字)
⚡ 30-Second TL;DR
有什麼變化
1,417 個過去預測問題與凍結 15M 文件語料庫,用於離線代理測試
為什麼重要
超越排行榜提供 AI 預測更深層分析,揭示特定策略弱點。引導代理在政治與商業等高風險領域改進。加速前沿模型更可靠推理的開發。
下一步行動
從 arXiv:2604.26106 下載 BTF-2,並在過去預測問題上基準測試你的預測代理。
誰應關注:Researchers & Academics
關鍵要點
- •1,417 個過去預測問題與凍結 15M 文件語料庫,用於離線代理測試
- •偵測 0.004 Brier 分數差異;區分研究與判斷優勢
- •透過預 mortem 與黑天鵝分析建構 0.011 Brier 更優預測器
- •前沿代理在領導者動機、計劃執行與機構流程上失敗
🧠 深度解析
AI-generated analysis for this event.
🔑 增強重點摘要
- •BTF-2 utilizes a 'frozen' corpus approach to solve the data contamination problem prevalent in LLM evaluation, ensuring that agents cannot access information published after the event date.
- •The benchmark introduces a novel 'Counterfactual Sensitivity' metric that measures how agents adjust their probability estimates when provided with specific, injected adversarial scenarios.
- •The research team behind BTF-2 identified that frontier models exhibit a 'confirmation bias' in forecasting, where they disproportionately weight evidence supporting their initial hypothesis over contradictory data.
📊 競品分析▸ Show
| Feature | BTF-2 | Metaculus (Platform) | Manifold Markets |
|---|---|---|---|
| Primary Use | Offline Agent Benchmarking | Human/Hybrid Forecasting | Prediction Markets |
| Data Source | Frozen 15M-doc Corpus | Real-time Web/Human Input | Real-time Market Data |
| Evaluation | Automated Brier Score | Crowd Consensus | Market Price |
| Pricing | Open Research/Free | Freemium | Transaction-based |
🛠️ 技術深入
- •Corpus Architecture: Uses a static, time-stamped snapshot of 15 million documents (news, policy papers, financial reports) indexed via a vector database to simulate a 'closed-world' information environment.
- •Evaluation Engine: Implements a multi-stage pipeline that separates 'Information Retrieval' (RAG performance) from 'Probabilistic Reasoning' (Brier score calculation) to isolate failure points.
- •Agent Prompting: Employs a chain-of-thought (CoT) framework specifically tuned for 'Pre-Mortem' analysis, forcing agents to generate three distinct failure scenarios before outputting a final probability distribution.
- •Sensitivity Analysis: Uses a perturbation-based testing method where key variables in the prompt are modified to check for non-linear shifts in the agent's confidence intervals.
🔮 前景展望AI analysis grounded in cited sources
Standardized forecasting benchmarks will become a mandatory component of AI safety evaluations.
Regulators are increasingly requiring quantitative evidence of an AI's ability to model complex, long-term strategic outcomes as a condition for deployment.
Future LLM architectures will incorporate dedicated 'probabilistic reasoning' modules.
The failure of current frontier models to handle institutional incentives suggests that general-purpose transformers require specialized layers for strategic game theory.
⏳ 時間線
2024-09
Initial release of BTF-1 focusing on basic geopolitical forecasting.
2025-06
Expansion of the frozen corpus to include specialized financial and scientific datasets.
2026-04
Official publication of BTF-2 with the 1,417-question dataset.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: ArXiv AI ↗