Chris Olah
Co-founder & Interpretability Lead, Anthropic
About
Chris Olah (b. 1992) is a Canadian machine-learning researcher and one of the pioneers of mechanistic interpretability — the effort to reverse-engineer neural networks by mapping the internal features and circuits that produce their behavior. A Thiel Fellow who left the University of Toronto after a year, he led interpretability research at Google Brain and OpenAI before co-founding Anthropic in 2021, where he leads the interpretability team. He co-founded Distill, a journal devoted to clear scientific communication, and his visual, essayistic explanations of neural networks shaped how a generation understands them. His work turns 'do we understand what we built?' from a rhetorical worry into an experimental science.
Key Contributions
- Pioneered mechanistic interpretability — reverse-engineering the features and circuits inside neural networks
- Co-founded Anthropic (2021) and leads its interpretability research
- Founded Distill, setting a new standard for clear, interactive scientific communication in ML
- Produced foundational work on feature visualization, circuits, and superposition in neural networks
- Named to the TIME100 AI list (2024) for advancing the science of understanding AI systems
Questions they sharpened View the streams
Videos & Interviews
Mechanistic Interpretability explained | Chris Olah and Lex Fridman
Olah explains what it means to find features and circuits inside a trained model.
View Details
Chris Olah - Looking Inside Neural Networks with Mechanistic Interpretability
Olah's 2023 Alignment Workshop talk on reverse-engineering the internals of neural networks.
View Details
The Dark Matter of AI [Mechanistic Interpretability]
Ask ChatGPT to forget a phrase and it will say it has, which is impossible — the phrase is in its context window — and press it and it will admit as much. We train assistants to be honest through examples; we have no direct access to whether they are. This is the interpretability problem, and the most promising handle on it is the sparse autoencoder, which extracts features that turn out to correspond to concepts a person can name.
View DetailsPapers & Publications
Connections
Dario Amodei
CollaboratedCEO & Co-founder, Anthropic
Their partnership predates both Anthropic and the LLM era: in 2016 they co-wrote 'Concrete Problems in AI Safety,' which turned vague fear of superintelligence into five researchable engineering failures. Five years later they left OpenAI together among Anthropic's seven founders, where Amodei runs the company that ships the models and Olah runs the team trying to read what is inside them. The bet that safety is an empirical science rather than a philosophical position runs from that paper straight through the lab.
arxiv.org · en.wikipedia.org
Amanda Askell
CollaboratedPhilosopher & AI Alignment Researcher
Both are on 'A General Language Assistant as a Laboratory for Alignment' (2021), Anthropic's first paper and the one that named the helpful-honest-harmless target Claude still aims at — Askell as first author, Olah near the end of the list. They approach the same model from opposite ends: she writes what it should be, he tries to read what it actually is. The distance between those two descriptions is roughly the whole of alignment research.
arxiv.org
Bret Victor
Influenced byInterface Designer & Computing Visionary
Victor coined 'explorable explanations' in 2011 for documents you can reach into and act on rather than only read. Six years later Olah and Shan Carter launched Distill on that premise and credited him in 'Research Debt', the essay arguing that fields accumulate an unpaid balance of interpretive labour — one explainer pays a fixed cost, every reader pays the understanding cost, and the multiplier is what makes distillation worth funding as research rather than as a favour. Distill is the clearest case of Victor's medium being taken up by a science that badly needed it, and Olah's later interpretability work keeps the habit: the circuits papers are things you scroll through and manipulate.
distill.pub · en.wikipedia.org
Geoffrey Hinton
KindredAI Pioneer & Researcher
Olah tells a joke on Hinton — that he has discovered how the brain works every year for fifty years — and is careful to add that he says it with deep respect, which is the right note for someone working the same seam from the other end. Hinton spent a career asserting that artificial networks and brains share principles; Olah's circuits programme put a version of that assertion where it could be checked. Its universality claim is that analogous features and circuits recur across models and tasks the way analogous organs recur across species. Where Hinton argued from the brain toward the network, Olah opens the network and looks.
youtube.com · distill.pub