🧠机器之心•較早收集於 11m
谷歌 DeepMind Aletheia 創 FirstProof 數學新紀錄

#math-ai#reasoning-agent#theorem-proving#benchmarkaletheiaaletheiadeepmindgemini-3firstproof
💡DeepMind AI cracks 6 real math research proofs—huge for automated theorem proving
⚡ 30-Second TL;DR
有什麼變化
自主解決 10 道未公開研究數學問題中的 6 道
為什麼重要
證明 AI 智能體能處理開放數學研究,彌合競賽解題與發現差距。加速自主定理證明。凸顯 DeepMind 在超人類推理領先。
下一步行動
Replicate Aletheia prompts from GitHub on your math agent for FirstProof problems.
誰應關注:Researchers & Academics
關鍵要點
- •自主解決 10 道未公開研究數學問題中的 6 道
- •FirstProof 最佳紀錄,超越 AI IMO 金牌
- •Gemini 3 Deep Think 驅動;論文與 GitHub 提示公開
- •專家審核證明,模擬真實數學研究標準
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 6 個來源。
🔑 增強重點摘要
- •Aletheia solved specific FirstProof problems 2, 5, 7, 8, 9, and 10, with expert disagreement only on problem 8[1][2].
- •Raw prompts and outputs for Aletheia are publicly available on GitHub at google-deepmind/superhuman/tree/main/aletheia[3].
- •Aletheia uses Google Search and web browsing as tools to prevent citation hallucinations and synthesize mathematical literature[5][6].
🛠️ 技術深入
- •Aletheia employs agentic scaffolding with iterative generation, verification, and revision using a natural language verifier to identify flaws[1][6].
- •Features two variants (Aletheia A and B) with best-of-2 submissions per problem, showing improved accuracy over December 2025 version via scaffolding and base model upgrades[2].
- •Integrates Gemini 3 Deep Think with inference-time scaling, achieving higher reasoning quality at lower compute (100x reduction from prior versions)[4][5].
🔮 前景展望AI analysis grounded in cited sources
Aletheia accelerates PhD-level math research by enabling autonomous paper generation.
AI resolves select open math problems at scale.
Aletheia autonomously solved 4 out of 700 Erdős conjectures and 63 technically correct solutions[5].
⏳ 時間線
2025-07
Gemini Deep Think achieves IMO gold-medal standard and 65.7% on IMO-ProofBench Advanced
2025-12
Early Aletheia version used for semi-autonomous Erdős problems
2026-01
Gemini Deep Think version achieves 95.1% on IMO-ProofBench Advanced with 100x compute reduction; Aletheia hits FutureMath Basic SOTA
2026-02-13
Aletheia submits best-of-2 solutions to FirstProof challenge
2026-02-24
arXiv paper released reporting Aletheia solving 6/10 FirstProof problems
📎 來源 (6)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- arXiv — 2602
- arXiv — 2602
- deeplearn.org — Aletheia Tackles Firstproof Autonomously
- atalupadhyay.wordpress.com — Aletheia Unveiled Googles Autonomous Mathematical Research AI
- marktechpost.com — Google Deepmind Introduces Aletheia the AI Agent Moving From Math Competitions to Fully Autonomous Professional Research Discoveries
- Google DeepMind — Accelerating Mathematical and Scientific Discovery with Gemini Deep Think
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 机器之心 ↗
每週 AI 簡報
每週一封,可隨時退訂。