🧠Stalecollected in 31m

Google outperforms OpenAI in recent math benchmark testing

Google outperforms OpenAI in recent math benchmark testing
PostLinkedIn
🧠Read original on The Neuron

💡Google's 9:1 lead in math benchmarks signals a major shift in the LLM competitive landscape.

⚡ 30-Second TL;DR

What Changed

Google's AI models demonstrate superior performance in complex mathematical problem-solving.

Why It Matters

This shift suggests that Google is regaining momentum in the LLM race, potentially forcing OpenAI to accelerate their next-generation model releases to maintain parity.

What To Do Next

Review the latest Google DeepMind research papers on mathematical reasoning to understand the architectural shifts driving these benchmark improvements.

Who should care:Researchers & Academics

Key Points

  • Google's AI models demonstrate superior performance in complex mathematical problem-solving.
  • The reported performance gap between Google and OpenAI is measured at a 9:1 ratio.
  • Mathematical reasoning remains a critical frontier for LLM development and competitive benchmarking.

🧠 Deep Insight

Web-grounded analysis with 14 cited sources.

🔑 Enhanced Key Takeaways

  • Google's mathematical agent Aletheia, powered by Gemini 3 Deep Think, autonomously solved 6 out of 10 questions in the challenging FirstProof competition, a benchmark designed to verify independent scientific research ability, outperforming OpenAI's internal model which solved 5 questions with human intervention.
  • OpenAI's general-purpose reasoning model recently disproved an 80-year-old conjecture related to Paul Erdős's planar unit distance problem, demonstrating a significant advance in AI reasoning for open mathematical problems, rather than relying on a model specifically trained for mathematics.
  • New benchmarks like SOOHAK are emerging to evaluate AI models on research-level mathematics and their ability to identify unsolvable problems, where Google's Gemini 3 Pro leads on solvable tasks but no model currently excels at recognizing problems without solutions.
  • Both Google DeepMind and OpenAI models achieved gold medal-level performance at the International Mathematical Olympiad (IMO) in July 2025, solving five out of six problems and signaling a breakthrough in AI's ability to tackle competition-level math.
  • The performance of top AI models is converging on traditional benchmarks like GPQA Diamond and MMLU, necessitating new evaluations such as Humanity's Last Exam (HLE) and ARC-AGI-2 to differentiate advanced reasoning capabilities.
📊 Competitor Analysis▸ Show
Feature/BenchmarkGoogle (Gemini 3.1 Pro / Deep Think / Aletheia)OpenAI (GPT-5.4 / o3 / o1)Anthropic (Claude Opus 4.6)DeepSeek (DeepSeek-R1)
MATH Benchmark95.1% (Gemini 3.1 Pro Preview)88.6% (GPT-5.4), 99.4% (GPT-5)N/AComparable to OpenAI-o1
AIME 2025N/A100% (GPT-5.4)N/AN/A
AIME 2024N/A74% (o1, single sample), 93% (o1, re-ranking 1000 samples)N/AN/A
FirstProof Challenge6/10 problems solved autonomously (Aletheia)5/10 problems solved with human intervention (internal model)N/AN/A
SOOHAK (Research-level)Leads at 30% on solvable problems (Gemini 3 Pro)N/AN/AN/A
GPQA Diamond94.3% (Gemini 3.1 Pro)94.4% (GPT-5.4), Exceeds human PhD-level accuracy (o1)89.6% (Claude Opus 4.6)N/A
ARC-AGI-284.6% (Gemini 3.1 Pro Deep Think)N/A38% (Claude Opus 4.6)N/A
HLE (Humanity's Last Exam)48.4% (Gemini 3.1 Pro Deep Think)41.6% (GPT-5.4)53.0% (Claude Opus 4.6)N/A
Pricing (Input/Output per 1M tokens)$2.00 / $12.00 (Gemini 3.1 Pro Preview)$2.50 / $15.00 (GPT-5.4)$5.00 / $25.00 (Claude Opus 4.6)N/A
Context Window1M tokens (Gemini 3.1 Pro)1M tokens (GPT-5.4)200K (1M beta) (Claude Opus 4.6)164K (DeepSeek-R1)
Key CapabilitiesAbstract reasoning, knowledge breadth, novel pattern recognition, agentic workflowsGeneral-purpose reasoning, mathematical proofs, formal logic, extended thinking timeNuanced analysis, complex code debuggingElite-level performance, RL-powered reasoning

🛠️ Technical Deep Dive

  • Google Gemini 3 Deep Think / Aletheia: This model is designed as a mathematical research agent capable of iteratively generating, verifying, and revising solutions for research-level math problems. It employs agentic reasoning workflows and has shown significant progress beyond Olympiad-level to PhD-level exercises.
  • OpenAI o3: A specialized reasoning model that utilizes extended 'thinking' time to enhance its problem-solving capabilities, particularly excelling in mathematical proofs, formal logic, and competition-style problems.
  • OpenAI General-Purpose Reasoning Model: The model that disproved the Erdős conjecture was a general-purpose reasoning model, not one specifically trained for mathematics, indicating advanced capabilities in applying broad reasoning to novel, open problems.
  • DeepSeek-R1: Features a Mixture-of-Experts (MoE) architecture with 671 billion total parameters and a 164K context length, powered by reinforcement learning (RL) for enhanced reasoning.
  • Qwen/QwQ-32B: A 32-billion parameter reasoning model incorporating advanced architectural elements such as Rotary Position Embeddings (RoPE), SwiGLU activation functions, RMSNorm, and Attention QKV bias, with 64 layers and a Grouped-Query Attention (GQA) architecture (40 Q attention heads, 8 for KV).

🔮 Future ImplicationsAI analysis grounded in cited sources

AI models will increasingly act as autonomous research partners in advanced mathematics and scientific discovery.
Recent breakthroughs, such as OpenAI disproving an 80-year-old conjecture and Google's Aletheia autonomously solving research-level problems, demonstrate AI's growing capacity for independent scientific contribution.
The focus of AI benchmarking will shift towards evaluating models on novel, unsolved, and even unsolvable problems to truly differentiate frontier capabilities.
Current benchmarks are saturating, and new evaluations like SOOHAK are being developed to test AI's ability to handle research-level math and recognize problems without solutions, pushing beyond memorization and curated datasets.
Human-AI collaboration in mathematics will evolve, with AI proposing novel constructions and approaches that human mathematicians then validate and refine.
The validation of OpenAI's Erdős problem solution involved human mathematicians improving the AI's initial proof, suggesting a future where AI generates possibilities for human experts to explore.

Timeline

2024-09
OpenAI's o1 model significantly improves AIME performance over GPT-4o.
2025-07
OpenAI and Google DeepMind models achieve gold medal-level performance at the International Mathematical Olympiad (IMO).
2026-02
Google DeepMind announces Gemini Deep Think's progress to PhD-level math and the Aletheia agent.
2026-02
Google's Aletheia outperforms OpenAI's internal model in the FirstProof challenge for independent scientific research ability.
2026-05
New SOOHAK benchmark released, testing AI on research-level math and unsolvable problems, with Google's Gemini 3 Pro leading on solvable tasks.
2026-05
OpenAI's general-purpose reasoning model disproves an 80-year-old conjecture related to Paul Erdős's planar unit distance problem.

📎 Sources (14)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. 36kr.com
  2. theguardian.com
  3. forbes.com
  4. eweek.com
  5. the-decoder.com
  6. deepmind.google
  7. mlq.ai
  8. humanprogress.org
  9. teamai.com
  10. apiyi.com
  11. pricepertoken.com
  12. siliconflow.com
  13. openai.com
  14. krater.ai
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Neuron