Human Tech Tree
Researched1992 · Digital Age (1990 – 2015)

Information / AI & Robotics

Learning by reward: TD-Gammon

Tesauro's TD-Gammon (1992) learned backgammon by playing itself, proving that reward-driven learning with neural nets works.

Open in the interactive tree →

Reinforcement learning lets a program learn from rewards instead of examples: Richard Sutton formalised temporal-difference learning in 1988 and Chris Watkins added Q-learning in 1989, building on animal-learning psychology. In 1992 Gerald Tesauro at IBM combined it with a neural network in TD-Gammon, which learned backgammon almost entirely by self-play and reached near-expert level. The same recipe, scaled up, later produced AlphaGo and the reward-trained reasoning models.

Prerequisites

Unlocks

Sources

More in AI & Robotics · Digital Age

All 39 points in AI & Robotics →

Open in the interactive tree →