Get Started
Research questionHow can we reliably explain what already-localized neural circuit components do?After a neural circuit has been localized, determining what its components do remains labor-intensive and difficult to standardize. Explanations also require validation to distinguish plausible interpretations from causally supported ones.
AI Agents
Evaluation & Benchmarks
Mechanistic Interpretability
Latest papersRecent research connected to this question, newest first.Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?Evidence comes from 84 semi-synthetic transformer circuits with 163 component annotations, evaluated across four language-model backbones, plus a case study of an arithmetic circuit in Llama-3-8B. The results indicate that validation failures, code execution errors, and unresolved hypotheses remain important limitations, and no backbone is uniformly best.research paper · Sep 2, 2026
Related questions
How can we measure whether transformer representations distinguish senses of the same word across contexts?How can causal transformers use task state discovered late in a long context to guide rereading?How can attention-head contributions be measured in prompt-injection classifiers across circuit and output scales?What can LLM–brain representational alignment establish about shared neural mechanisms—and what can’t it?
Home
Topics
Search
Library