lecture December 23, 2024 24:09 YouTube

The Dark Matter of AI [Mechanistic Interpretability]

Welch Labs

AIInterpretabilityAlignment

Ask ChatGPT to forget a phrase and it will say it has, which is impossible — the phrase is in its context window — and press it and it will admit as much. We train assistants to be honest through examples; we have no direct access to whether they are. This is the interpretability problem, and the most promising handle on it is the sparse autoencoder, which extracts features that turn out to correspond to concepts a person can name.

Theme
Language
Support
© funclosure 2025