Research questionHow can we test whether language models causally use scientific mechanisms instead of answer-correlated shortcuts?Correct scientific answers do not by themselves show that a model represents the governing mechanism or uses it to decide. Numerical and lexical patterns can support accurate outputs without mechanistic understanding.