Mechanistic Interpretability: A Practical Field Guide for Applied AI Teams
- Published
- Author
Dhiru Jadhav
A practical introduction to looking inside language models, finding the features and circuits behind a behavior, and testing whether an explanation is causal.
Read more