#047 · Gods of AIEpicLocked in

Paul Christiano

The RLHF Paper

Learning from human preferences

Rating

87Impact score
Research94
Systems built80
Openness88
Influence91
FromUSA
Era2017–present

Domains

AlignmentPreference learningAI safety

Scouting report

Lead author of deep reinforcement learning from human preferences, the technique that became RLHF and turned raw language models into things that follow instructions — the single largest reason chat assistants work at all. Later founded the Alignment Research Center and moved to the US AI Safety Institute.

Curator’s note

Wrote the alignment paper that turned out to also be the product paper.