Open Nemotron Recipe Reaches IMO Gold Threshold

An open model reaches the IMO gold threshold using only natural-language inference.
30-Second TL;DR
What Changed
The system uses supervised fine-tuning and reinforcement learning.
Why It Matters
The work offers an unusually reproducible recipe for improving mathematical reasoning in open models. It could accelerate research into inference-time search, self-verification, and proof refinement without relying on proprietary tools.
What To Do Next
Clone the released Nemotron checkpoints and benchmark their proof-verification loop against your current reasoning pipeline.
Key Points
- •The system uses supervised fine-tuning and reinforcement learning.
- •Three Nemotron 3 Ultra checkpoints generate, verify, and refine proofs.
- •The pipeline uses no formal prover, external tools, or internet access.
- •Researchers released checkpoints, data, code, solutions, and a 200-problem benchmark.
Deep Insight
Background and context from public sources — not the original article. 10 sources cited.
Enhanced Key Takeaways
- •The technical paper (arXiv:2609.10712), titled 'An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics', was authored by NVIDIA researchers Ivan Moshkov, Stephen Ge, George Armstrong, Wei Du, Sadegh Mahdavi, and Igor Gitman.
- •The system's 30 out of 42 evaluation score surpassed the IMO 2026 gold-medal threshold of 29 points by a single point under natural-language grading conditions.
- •The specific domain-adapted specialist checkpoints deployed are designated as Nemotron-3-Labs-Ultra-Math-SFT and Nemotron-3-Labs-Ultra-Math-RL, published under the OpenMDW-1.1 license.
- •The released 200-problem evaluation suite is officially designated as 'Nemotron-IMO-Bench', providing novel olympiad-grade problems to evaluate post-training reasoning pipelines.
- •This achievement expands NVIDIA's 2026 olympiad reasoning lineup, following the 30B MoE Nemotron-Cascade 2 (IMO/IOI 2025) and the competitive coding system Nemotron-3-Ultra-CC (IOI 2026).
Competitor Analysis
- Primary Developer
- NVIDIA
- Modality & Approach
- Pure natural-language reasoning (two-stage test-time compute, no provers)
- Access Model
- Open weights & open code (OpenMDW-1.1)
- Benchmark Performance
- 30/42 at IMO 2026 (Gold)
- Primary Developer
- DeepSeek
- Modality & Approach
- 671B MoE natural-language test-time reasoning
- Access Model
- Open weights
- Benchmark Performance
- Gold threshold at IMO-level mathematics
- Primary Developer
- Google DeepMind
- Modality & Approach
- Hybrid formal language (Lean) & natural-language search
- Access Model
- Proprietary
- Benchmark Performance
- Gold/Silver medal equivalent at IMO 2024–2025
- Primary Developer
- OpenAI
- Modality & Approach
- Test-time reasoning and reinforcement learning (natural language)
- Access Model
- Proprietary API
- Benchmark Performance
- Competitive Olympiad-level performance
| System / Model | Primary Developer | Modality & Approach | Access Model | Benchmark Performance |
|---|---|---|---|---|
| Nemotron 3 Ultra IMO Pipeline | NVIDIA | Pure natural-language reasoning (two-stage test-time compute, no provers) | Open weights & open code (OpenMDW-1.1) | 30/42 at IMO 2026 (Gold) |
| DeepSeek-V3.2-Speciale | DeepSeek | 671B MoE natural-language test-time reasoning | Open weights | Gold threshold at IMO-level mathematics |
| AlphaProof / Gemini Deep Think | Google DeepMind | Hybrid formal language (Lean) & natural-language search | Proprietary | Gold/Silver medal equivalent at IMO 2024–2025 |
| OpenAI Reasoning Series (o-series) | OpenAI | Test-time reasoning and reinforcement learning (natural language) | Proprietary API | Competitive Olympiad-level performance |
Technical Deep Dive
- Core Backbone & Specialized Checkpoints: Built on NVIDIA's Nemotron 3 Ultra architecture utilizing a three-checkpoint setup: the general-availability foundation model, a supervised fine-tuning specialist (
Nemotron-3-Labs-Ultra-Math-SFT), and a reinforcement-learning specialist (Nemotron-3-Labs-Ultra-Math-RL). - Two-Stage Test-Time Compute (TTC): Operates via an iterative generation-critique-refine loop in Stage 1 where specialist models debate and revise candidate proofs, followed by Stage 2 independent high-compute verification and ranking to select the final proof transcript.
- Unassisted Natural-Language Execution: Generates fully informal, natural-language mathematical proofs without reliance on formal interactive theorem provers (e.g., Lean, Isabelle), symbolic computer algebra engines, or live web retrieval.
- Artifact Release Scope: Distributed via the
nvidia/nemotron-labs-imo-2026repository, containing the full training corpus, inference scripts, solution transcripts, and theNemotron-IMO-Bench200-problem benchmark under OpenMDW-1.1.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2026-03NVIDIA releases Nemotron-Cascade 2 (30B MoE) achieving gold-medal performance on IMO 2025 and IOI 2025 problems
- 2026-09NVIDIA introduces Nemotron-3-Ultra-CC tailored for IOI 2026 competitive programming
- 2026-09NVIDIA researchers publish arXiv:2609.10712 unveiling the Nemotron 3 Ultra IMO 2026 gold-level recipe
- 2026-09NVIDIA open-sources Nemotron-3-Labs-Ultra checkpoints and the Nemotron-IMO-Bench under OpenMDW-1.1
Sources (10)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.