
I’m Tom Conerly a researcher on the interpretability team at Anthropic. I’ll give a whirlwind tour of our latest paper “Circuit Tracing: Revealing Computational Graphs in Language Models”. We’ll go over how Claude plans rhymes in poetry, how a simple jailbreak works, and how Claude shares computation across languages. There will be plenty of concrete examples and time for questions.
Paper links:
https://transformer-circuits.pub/2025/attribution-graphs/methods.html
https://transformer-circuits.pub/2025/attribution-graphs/biology.html