
What is our current understanding of the inner workings of LLMs? I’ll go over work I did as part of the Anthropic interpretability team where we found millions of human interpretable features in Claude 3 Sonnet.
We found that most neurons are polysemantic meaning they activate on many different human concepts. We used “dictionary learning”, borrowed from classical machine learning, to find monosemantic features from polysemantic neurons.
https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html