Researched
Learning from Human Feedback
Christiano et al. (2017) teach models from human preference comparisons; OpenAI's InstructGPT (2022) uses it to make language models follow instructions.
Open in the interactive tree →People rank model answers, a reward model learns the ranking, and reinforcement learning tunes the model. It turned raw language models into assistants like ChatGPT and is the main practical alignment method.
Prerequisites
Unlocks
- Large Language Models2022Human-feedback training turned language models into usable assistants
- AI Alignment & InterpretabilityopenLearning from human feedback is the main practical alignment method