Reasoning Models
Models that think at length before answering and are trained with reward learning crack olympiad math and hard programming tasks.
Open in the interactive tree →OpenAI o1 (September 2024) and DeepSeek-R1 (January 2025, open weights) train models with reinforcement learning on verifiable tasks (math, code), so they produce longer chains of thought and correct themselves. More compute at answer time yields better results: a second scaling axis next to training.
As of October 2026
Reasoning is standard in all top models in 2026. In July 2025 Google DeepMind’s Gemini Deep Think earned an officially graded gold medal at the International Mathematical Olympiad (35 of 42 points), and in July 2026 Huawei and Xiaohongshu reported perfect 42/42 scores for their systems. On the programming benchmark SWE-bench Verified, the top score rose from 60% to nearly 100% within a year, per the Stanford AI Index 2026.
Open steps
- Reward learning without an answer key High AI leverageReasoning training works where answers can be checked; open-ended science, law or writing need reward signals that cannot be gamed.
- Less overthinking, cheaper answers High AI leverageReasoning models spend thousands of tokens even on easy questions; adaptive thinking length would cut cost and delay without losing accuracy.
- Trustworthy chains of thought Medium AI leverageThe written reasoning may not match what drives the answer, and training can make it unreadable; keeping it faithful would allow monitoring for misbehaviour.
Where AI could help
Medium AI leverage. Models can generate and grade training problems wherever answers can be checked automatically; open-ended domains lack cheap checkers.
- Generate large sets of verifiable problems and graders for reinforcement learning
- Use formal proof checkers as exact rewards for math and code
- Distill long reasoning into cheaper models
- Analyze failure cases and design curricula with agents
Shown so far
- In January 2025 the DeepSeek-R1 paper (later in Nature) showed reasoning abilities such as self-verification emerging from reinforcement learning without human-labeled reasoning examples. source
- In July 2024 DeepMind's AlphaProof and AlphaGeometry 2 solved 4 of 6 IMO problems with machine-checkable formal proofs (according to the company; method published in Nature, November 2025). source
Prerequisites
- AlphaGo2016
- Large Language Models2022
- Open-Weight Language Models2023-2026DeepSeek-R1 (January 2025) released an RL-trained reasoning model with open weights
Unlocks
- AI Agents2025
- AI Wins Math Olympiad Gold2025AI olympiad gold came from reasoning models that think at length before answering
- Artificial General Intelligenceopen
- Reliable, Honest AIopen