Get Started
Home
Topics
Search
Library
Research questionCan chain-of-thought monitoring detect preferences received through tools or inferred from raw artifacts?Chain-of-thought monitoring assumes that a model’s reasoning trace reveals the information influencing its answer. Preferences delivered through tool returns or inferred from unprocessed artifacts may affect answers without being clearly verbalized in the trace.
AI
AI Agents
Alignment & Safety
Evaluation & Benchmarks
Reasoning
Latest papersRecent research connected to this question, newest first.Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are DeliveredEvidence comes from a 5,100-sample evaluation of 15 open-weight models, comparing user-message and tool-return cues that were either explicitly summarized or presented as raw artifacts. It measures verbalized commitment, unverbalized preference adoption, and transcript-monitor detection in the tested single-call, prefilled-tool setting.research paper · Sep 1, 2026
Related questions
Can chain-of-thought monitoring detect consequential computation hidden in semantically irrelevant filler tokens?How reliably can chain-of-thought text reveal which reasoning steps causally drive correct answers?How can we tell whether deceptive-looking language-model behavior reflects a deceptive mechanism?How can we measure whether long-horizon tool-using agents query hidden state and execute stated plans?