

Tri Dao
Author of FlashAttention
Making attention fit in memory, and co-creating Mamba
Rating
89Impact score
Domains
Efficient attentionState space modelsGPU kernels
Scouting report
Wrote FlashAttention, an IO-aware exact attention kernel that cut memory use enough to make long contexts practical, and released it openly — it is now in essentially every training stack in the world. Co-created the Mamba state space architecture with Albert Gu, also open. Chief scientist at Together AI and a Princeton professor.
Curator’s note
If your context window is long, it is partly because of his CUDA.
Portrait
Painted from a photograph of Tri Dao by Tri Dao, supplied by the subject.
Sources
- 01FlashAttention (2022) arxiv.org
- 02flash-attention on GitHub github.com