

Christoph Schuhmann
Founder of LAION
The open datasets the image models were trained on
Rating
84Impact score
Domains
Open datasetsMultimodal dataCommunity research
Scouting report
A schoolteacher who organised a volunteer effort to assemble LAION-400M and LAION-5B, the open image-text datasets that Stable Diffusion and OpenCLIP were trained on. Before LAION, the data behind multimodal models was private by default; afterwards, anyone could inspect it — including the researchers who found serious problems in it.
Curator’s note
Made the training data auditable, which is how its flaws got found.
Portrait
Painted from a photograph of Christoph Schuhmann by Christoph Schuhmann, supplied by the subject.
Sources
- 01LAION-5B (2022) arxiv.org
- 02LAION laion.ai