Researched
AlphaGo
Google’s AlphaGo beats world-class Go player Lee Sedol 4-1, in a game long thought out of reach for decades.
Open in the interactive tree →AlphaGo combined deep networks with tree search and self-play (reinforcement learning). In March 2016 it beat Lee Sedol 4-1; AlphaGo Zero (2017) learned without human games and surpassed it. The same family of methods led to AlphaFold (protein structures, 2024 Nobel Prize in Chemistry) and to reasoning models.
Prerequisites
- Monte Carlo Methods1949AlphaGo used Monte Carlo tree search to choose moves
- Learning by reward: TD-Gammon1992Self-play reinforcement learning was proven with TD-Gammon
- Deep Blue Beats Kasparov1997
- Deep Learning2012
Unlocks
- Reasoning Models2024