Chris Olah
Co-founder & Interpretability Lead, Anthropic
關於
Chris Olah (b. 1992) is a Canadian machine-learning researcher and one of the pioneers of mechanistic interpretability — the effort to reverse-engineer neural networks by mapping the internal features and circuits that produce their behavior. A Thiel Fellow who left the University of Toronto after a year, he led interpretability research at Google Brain and OpenAI before co-founding Anthropic in 2021, where he leads the interpretability team. He co-founded Distill, a journal devoted to clear scientific communication, and his visual, essayistic explanations of neural networks shaped how a generation understands them. His work turns 'do we understand what we built?' from a rhetorical worry into an experimental science.
主要貢獻
- Pioneered mechanistic interpretability — reverse-engineering the features and circuits inside neural networks
- Co-founded Anthropic (2021) and leads its interpretability research
- Founded Distill, setting a new standard for clear, interactive scientific communication in ML
- Produced foundational work on feature visualization, circuits, and superposition in neural networks
- Named to the TIME100 AI list (2024) for advancing the science of understanding AI systems
影片與訪談
Mechanistic Interpretability explained | Chris Olah and Lex Fridman
Olah explains what it means to find features and circuits inside a trained model.
View Details
Chris Olah - Looking Inside Neural Networks with Mechanistic Interpretability
Olah's 2023 Alignment Workshop talk on reverse-engineering the internals of neural networks.
View Details