Get Started
Home
Topics
Search
Library
Research questionHow can we tell whether deceptive-looking language-model behavior reflects a deceptive mechanism?A language model can produce misleading outputs for reasons that do not involve a stable deceptive preference or objective. This makes it difficult to infer what mechanism, if any, generated the behavior.
AI
Alignment & Safety
Evaluation & Benchmarks
Mechanistic Interpretability
Latest papersRecent research connected to this question, newest first.From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception ResearchThe framework is tested on two open-weight model families using controlled guessing-game and stock-trading experiments. It distinguishes several causal sources of apparent deception and reports interventions showing that recipient information can affect deceptive preference, while noting that evidence for a deceptive mechanism does not establish model agency.research paper · Sep 3, 2026
Related questions
How can we test whether language models causally use scientific mechanisms instead of answer-correlated shortcuts?Do more capable language models exhibit stable task preferences that conflict with helpful, honest behavior?How can we distinguish decodable logical validity from reasoning that actually drives a language model’s answers?How can evaluators distinguish missing knowledge from miscalibrated outputs in language models?