The Dark Matter of AI [Mechanistic Interpretability]
AIInterpretabilityAlignment
Ask ChatGPT to forget a phrase and it will say it has, which is impossible — the phrase is in its context window — and press it and it will admit as much. We train assistants to be honest through examples; we have no direct access to whether they are. This is the interpretability problem, and the most promising handle on it is the sparse autoencoder, which extracts features that turn out to correspond to concepts a person can name.