

Jan Leike
The Alignment Lead
InstructGPT, superalignment, and resigning over it
Rating
84Impact score
Domains
AlignmentRLHFScalable oversight
Scouting report
Co-led the work that turned RLHF into InstructGPT and then ChatGPT, and headed OpenAI’s superalignment effort. Resigned publicly in 2024 saying safety culture had lost out to shipping, which is a rare thing for someone at that level to say on the record, and moved to Anthropic to continue the work.
Curator’s note
Built the alignment method, then said out loud when it was being sidelined.
Portrait
Painted from a photograph of Jan Leike by Jan Leike, supplied by the subject.
Sources
- 01Training language models to follow instructions (2022) arxiv.org
- 02Jan Leike en.wikipedia.org