Human Tech Tree
Researched2017 · Present (2015 – Oct 2026)

Information / AI & Robotics

Learning from Human Feedback

Christiano et al. (2017) teach models from human preference comparisons; OpenAI's InstructGPT (2022) uses it to make language models follow instructions.

Open in the interactive tree →

People rank model answers, a reward model learns the ranking, and reinforcement learning tunes the model. It turned raw language models into assistants like ChatGPT and is the main practical alignment method.

Prerequisites

Unlocks

Sources

More in AI & Robotics · Present

All 39 points in AI & Robotics →

Open in the interactive tree →