

Georgi Gerganov
Author of llama.cpp
Putting large models on ordinary hardware
Rating
91Impact score
Domains
InferenceQuantisationLocal models
Scouting report
Wrote ggml, whisper.cpp and llama.cpp — dependency-free C implementations that run large models on laptops, phones and Raspberry Pis. The GGUF format and its quantisation schemes are how most people actually run a model locally. Arguably did more for practical access to open models than any lab release.
Curator’s note
One person, plain C, and suddenly the models ran on your own machine.
Portrait
Painted from a photograph of Georgi Gerganov by Georgi Gerganov, supplied by the subject.
Sources
- 01llama.cpp github.com
- 02whisper.cpp github.com