AI Wins Math Olympiad Gold
In July 2025 AI systems reach gold level at the International Mathematical Olympiad (35/42); in July 2026 Huawei and Xiaohongshu claim perfect 42/42 scores under IMO judging (company claims).
Open in the interactive tree →The IMO is the benchmark for creative proof-writing. Google DeepMind's Gemini Deep Think solved five of six problems at IMO 2025 (35/42 points, gold) in natural language and was graded officially; OpenAI reported the same score from its own grading. In 2024 AlphaProof and AlphaGeometry 2 had reached silver-medal level (28/42) with formal Lean proofs.
As of October 2026
At IMO 2026 (Shanghai, papers on 15-16 July, 666 contestants, 7 with full marks), Huawei's Celia and Xiaohongshu's dots-note-3.0 each scored 42/42 under the official IMO judging process, the first perfect scores by language models under official IMO judging, according to AFP. An investor also reported that models from OpenAI, Anthropic, Axiom Math and Moonshot AI scored 42/42 in their own tests; these claims are unconfirmed. Olympiad problems are thus effectively solved, and attention has moved to open research problems.
Open steps
- Research-level problem sets High AI leverageBuild held-out sets of genuine research questions with checkable answers or Lean statements, since olympiad tests are saturated at 42/42 and no longer separate systems.
- Reliable grading of written proofs High AI leverageMake automatic judges of natural-language proofs dependable on long, novel arguments, so a false pass is rare and no human has to read every attempt.
- From olympiad tricks to new lemmas Medium AI leverageTest whether systems that solve short olympiad problems can prove multi-step lemmas that need new definitions, using unproved lemmas in published papers as the testbed.
- Gold-level proofs written in Lean High AI leverageReach full-mark olympiad performance with proofs that a Lean checker accepts end to end, so correctness no longer depends on human or model graders.
Where AI could help
High AI leverage. AI is the subject here: the step from olympiad to research-level proofs is what labs scale with compute and Lean checking.
- Replace saturated olympiad tests with unpublished research-level problems that have checkable answers
- Pair every proof attempt with Lean checking so results are verifiable without reading pages
- Run many agents in parallel on lemma search and literature for human mathematicians
Shown so far
- In July 2026 Huawei and Xiaohongshu said their models each scored 42/42 at the IMO under the official grading process, the first perfect scores by language models under official judging, according to AFP and the companies. source
- In November 2025 DeepMind published AlphaProof in Nature, a Lean-based reinforcement-learning agent that, with AlphaGeometry 2, reached silver-medal level at the 2024 Olympiad. source
Prerequisites
- Formal Proofs (Lean, Rocq)2005-2026
- Large Language Models2022
- Reasoning Models2024AI olympiad gold came from reasoning models that think at length before answering