Get Started
Home
Topics
Search
Library
Research questionHow can we test whether language models causally use scientific mechanisms instead of answer-correlated shortcuts?Correct scientific answers do not by themselves show that a model represents the governing mechanism or uses it to decide. Numerical and lexical patterns can support accurate outputs without mechanistic understanding.
AI
Evaluation & Benchmarks
Mechanistic Interpretability
Reasoning
Research Paper
Latest papersRecent research connected to this question, newest first.Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language ModelThe evidence concerns three open-weight Gemma 4 models answering materials-science questions. It examines mechanism representations and causal use through held-out descriptions, counterfactual laws, hidden-state comparisons, and state interventions; apparent organization in absolute states was confounded by numerical comparison, while controlled state changes provided stronger evidence of physical relationships.research paper · Sep 3, 2026
Related questions
How can language models produce reliable educational answers with reasoning that is verifiable?How can we tell whether deceptive-looking language-model behavior reflects a deceptive mechanism?How can we compare language models’ conditional behavior and predict the effects of prompt changes?How can evaluators distinguish missing knowledge from miscalibrated outputs in language models?