Research questionHow can we distinguish decodable logical validity from reasoning that actually drives a language model’s answers?A logical-validity signal can be readily decoded from hidden states even when the model answers incorrectly. The central difficulty is determining whether that signal is expressed in behavior and causally influences the answer.