Research questionHow can we tell whether deceptive-looking language-model behavior reflects a deceptive mechanism?A language model can produce misleading outputs for reasons that do not involve a stable deceptive preference or objective. This makes it difficult to infer what mechanism, if any, generated the behavior.