Researched
Learning by reward: TD-Gammon
Tesauro's TD-Gammon (1992) learned backgammon by playing itself, proving that reward-driven learning with neural nets works.
Open in the interactive tree →Reinforcement learning lets a program learn from rewards instead of examples: Richard Sutton formalised temporal-difference learning in 1988 and Chris Watkins added Q-learning in 1989, building on animal-learning psychology. In 1992 Gerald Tesauro at IBM combined it with a neural network in TD-Gammon, which learned backgammon almost entirely by self-play and reached near-expert level. The same recipe, scaled up, later produced AlphaGo and the reward-trained reasoning models.
Prerequisites
- Freud, Pavlov, behaviorism1900
- Markov Chains1906Reinforcement learning is built on Markov decision processes
- Operant conditioning1938Reinforcement learning formalises Skinner's idea that behaviour is shaped by reward
- Backpropagation1986
Unlocks
- Dopamine and reward prediction1997
- AlphaGo2016Self-play reinforcement learning was proven with TD-Gammon
- Learning from Human Feedback2017