

Paul Christiano
The RLHF Paper
Learning from human preferences
Rating
87Impact score
Domains
AlignmentPreference learningAI safety
Scouting report
Lead author of deep reinforcement learning from human preferences, the technique that became RLHF and turned raw language models into things that follow instructions — the single largest reason chat assistants work at all. Later founded the Alignment Research Center and moved to the US AI Safety Institute.
Curator’s note
Wrote the alignment paper that turned out to also be the product paper.
Sources
- 01Deep Reinforcement Learning from Human Preferences (2017) arxiv.org
- 02Alignment Research Center alignment.org