🧠較早收集於 11m

谷歌 DeepMind Aletheia 創 FirstProof 數學新紀錄

谷歌 DeepMind Aletheia 創 FirstProof 數學新紀錄
PostLinkedIn
🧠閱讀原文: 机器之心
#math-ai#reasoning-agent#theorem-proving#benchmarkaletheiaaletheiadeepmindgemini-3firstproof

💡DeepMind AI cracks 6 real math research proofs—huge for automated theorem proving

⚡ 30-Second TL;DR

有什麼變化

自主解決 10 道未公開研究數學問題中的 6 道

為什麼重要

證明 AI 智能體能處理開放數學研究,彌合競賽解題與發現差距。加速自主定理證明。凸顯 DeepMind 在超人類推理領先。

下一步行動

Replicate Aletheia prompts from GitHub on your math agent for FirstProof problems.

誰應關注:Researchers & Academics

關鍵要點

  • 自主解決 10 道未公開研究數學問題中的 6 道
  • FirstProof 最佳紀錄,超越 AI IMO 金牌
  • Gemini 3 Deep Think 驅動;論文與 GitHub 提示公開
  • 專家審核證明,模擬真實數學研究標準

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 6 個來源。

🔑 增強重點摘要

  • Aletheia solved specific FirstProof problems 2, 5, 7, 8, 9, and 10, with expert disagreement only on problem 8[1][2].
  • Raw prompts and outputs for Aletheia are publicly available on GitHub at google-deepmind/superhuman/tree/main/aletheia[3].
  • Aletheia uses Google Search and web browsing as tools to prevent citation hallucinations and synthesize mathematical literature[5][6].

🛠️ 技術深入

  • Aletheia employs agentic scaffolding with iterative generation, verification, and revision using a natural language verifier to identify flaws[1][6].
  • Features two variants (Aletheia A and B) with best-of-2 submissions per problem, showing improved accuracy over December 2025 version via scaffolding and base model upgrades[2].
  • Integrates Gemini 3 Deep Think with inference-time scaling, achieving higher reasoning quality at lower compute (100x reduction from prior versions)[4][5].

🔮 前景展望AI analysis grounded in cited sources

Aletheia accelerates PhD-level math research by enabling autonomous paper generation.
It produced the fully autonomous Feng26 paper on arithmetic geometry eigenweights without human mathematical intervention[4][5].
AI resolves select open math problems at scale.
Aletheia autonomously solved 4 out of 700 Erdős conjectures and 63 technically correct solutions[5].

時間線

2025-07
Gemini Deep Think achieves IMO gold-medal standard and 65.7% on IMO-ProofBench Advanced
2025-12
Early Aletheia version used for semi-autonomous Erdős problems
2026-01
Gemini Deep Think version achieves 95.1% on IMO-ProofBench Advanced with 100x compute reduction; Aletheia hits FutureMath Basic SOTA
2026-02-13
Aletheia submits best-of-2 solutions to FirstProof challenge
2026-02-24
arXiv paper released reporting Aletheia solving 6/10 FirstProof problems
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 机器之心

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週 AI 簡報

每週一封,可隨時退訂。