Google outperforms OpenAI in recent math benchmark testing

💡Google's 9:1 lead in math benchmarks signals a major shift in the LLM competitive landscape.
⚡ 30-Second TL;DR
What Changed
Google's AI models demonstrate superior performance in complex mathematical problem-solving.
Why It Matters
This shift suggests that Google is regaining momentum in the LLM race, potentially forcing OpenAI to accelerate their next-generation model releases to maintain parity.
What To Do Next
Review the latest Google DeepMind research papers on mathematical reasoning to understand the architectural shifts driving these benchmark improvements.
Key Points
- •Google's AI models demonstrate superior performance in complex mathematical problem-solving.
- •The reported performance gap between Google and OpenAI is measured at a 9:1 ratio.
- •Mathematical reasoning remains a critical frontier for LLM development and competitive benchmarking.
🧠 Deep Insight
Web-grounded analysis with 14 cited sources.
🔑 Enhanced Key Takeaways
- •Google's mathematical agent Aletheia, powered by Gemini 3 Deep Think, autonomously solved 6 out of 10 questions in the challenging FirstProof competition, a benchmark designed to verify independent scientific research ability, outperforming OpenAI's internal model which solved 5 questions with human intervention.
- •OpenAI's general-purpose reasoning model recently disproved an 80-year-old conjecture related to Paul Erdős's planar unit distance problem, demonstrating a significant advance in AI reasoning for open mathematical problems, rather than relying on a model specifically trained for mathematics.
- •New benchmarks like SOOHAK are emerging to evaluate AI models on research-level mathematics and their ability to identify unsolvable problems, where Google's Gemini 3 Pro leads on solvable tasks but no model currently excels at recognizing problems without solutions.
- •Both Google DeepMind and OpenAI models achieved gold medal-level performance at the International Mathematical Olympiad (IMO) in July 2025, solving five out of six problems and signaling a breakthrough in AI's ability to tackle competition-level math.
- •The performance of top AI models is converging on traditional benchmarks like GPQA Diamond and MMLU, necessitating new evaluations such as Humanity's Last Exam (HLE) and ARC-AGI-2 to differentiate advanced reasoning capabilities.
📊 Competitor Analysis▸ Show
| Feature/Benchmark | Google (Gemini 3.1 Pro / Deep Think / Aletheia) | OpenAI (GPT-5.4 / o3 / o1) | Anthropic (Claude Opus 4.6) | DeepSeek (DeepSeek-R1) |
|---|---|---|---|---|
| MATH Benchmark | 95.1% (Gemini 3.1 Pro Preview) | 88.6% (GPT-5.4), 99.4% (GPT-5) | N/A | Comparable to OpenAI-o1 |
| AIME 2025 | N/A | 100% (GPT-5.4) | N/A | N/A |
| AIME 2024 | N/A | 74% (o1, single sample), 93% (o1, re-ranking 1000 samples) | N/A | N/A |
| FirstProof Challenge | 6/10 problems solved autonomously (Aletheia) | 5/10 problems solved with human intervention (internal model) | N/A | N/A |
| SOOHAK (Research-level) | Leads at 30% on solvable problems (Gemini 3 Pro) | N/A | N/A | N/A |
| GPQA Diamond | 94.3% (Gemini 3.1 Pro) | 94.4% (GPT-5.4), Exceeds human PhD-level accuracy (o1) | 89.6% (Claude Opus 4.6) | N/A |
| ARC-AGI-2 | 84.6% (Gemini 3.1 Pro Deep Think) | N/A | 38% (Claude Opus 4.6) | N/A |
| HLE (Humanity's Last Exam) | 48.4% (Gemini 3.1 Pro Deep Think) | 41.6% (GPT-5.4) | 53.0% (Claude Opus 4.6) | N/A |
| Pricing (Input/Output per 1M tokens) | $2.00 / $12.00 (Gemini 3.1 Pro Preview) | $2.50 / $15.00 (GPT-5.4) | $5.00 / $25.00 (Claude Opus 4.6) | N/A |
| Context Window | 1M tokens (Gemini 3.1 Pro) | 1M tokens (GPT-5.4) | 200K (1M beta) (Claude Opus 4.6) | 164K (DeepSeek-R1) |
| Key Capabilities | Abstract reasoning, knowledge breadth, novel pattern recognition, agentic workflows | General-purpose reasoning, mathematical proofs, formal logic, extended thinking time | Nuanced analysis, complex code debugging | Elite-level performance, RL-powered reasoning |
🛠️ Technical Deep Dive
- Google Gemini 3 Deep Think / Aletheia: This model is designed as a mathematical research agent capable of iteratively generating, verifying, and revising solutions for research-level math problems. It employs agentic reasoning workflows and has shown significant progress beyond Olympiad-level to PhD-level exercises.
- OpenAI o3: A specialized reasoning model that utilizes extended 'thinking' time to enhance its problem-solving capabilities, particularly excelling in mathematical proofs, formal logic, and competition-style problems.
- OpenAI General-Purpose Reasoning Model: The model that disproved the Erdős conjecture was a general-purpose reasoning model, not one specifically trained for mathematics, indicating advanced capabilities in applying broad reasoning to novel, open problems.
- DeepSeek-R1: Features a Mixture-of-Experts (MoE) architecture with 671 billion total parameters and a 164K context length, powered by reinforcement learning (RL) for enhanced reasoning.
- Qwen/QwQ-32B: A 32-billion parameter reasoning model incorporating advanced architectural elements such as Rotary Position Embeddings (RoPE), SwiGLU activation functions, RMSNorm, and Attention QKV bias, with 64 layers and a Grouped-Query Attention (GQA) architecture (40 Q attention heads, 8 for KV).
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (14)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Neuron ↗

