Get Started
Research questionHow can we tell whether internal estimates guide effective actions, rather than merely predict effects accurately?An estimate can be accurate on average yet point to a harmful or ineffective action when reachable intervention directions and differing costs shape the outcome. The quality of the selected action is therefore distinct from the accuracy of the estimate.
Alignment & Safety
Evaluation & Benchmarks
Mechanistic Interpretability
Latest papersRecent research connected to this question, newest first.ObserverBench: Testing Mechanistic Estimates for Intervention and ControlEvidence covers circuit-intervention studies on GPT-2-small and Qwen2.5-7B, safety-triage tasks, and APPS tasks on Qwen3.5-9B, with monitor panels spanning Qwen2.5-7B and Gemma-2-9B-it. Pairwise observers can predict unseen effects more accurately without choosing better actions, while action-loss-trained observers choose lower-loss actions; AUROC rankings can differ from deployment loss, the best information source varies by model, and sparse SAE readouts trail dense controls on reported Qwen panels. These findings are limited to the reported models, tasks, and task contracts.research paper · Sep 2, 2026
Related questions
How can world models guide safe intervention in embodied systems when likely futures omit consequences and uncertainty?How can evaluators distinguish missing knowledge from miscalibrated outputs in language models?How can we measure whether long-horizon tool-using agents query hidden state and execute stated plans?How can we distinguish genuine agentic progress from interface expansion, persistence, and environmental coupling when delegating authority?
Home
Topics
Search
Library