Get Started
Home
Topics
Search
Library
Research questionCan chain-of-thought monitoring detect consequential computation hidden in semantically irrelevant filler tokens?Language models may gain task performance from semantically irrelevant filler tokens without making the relevant computation interpretable in their visible reasoning. This complicates the use of chain-of-thought as evidence of what a model has computed.
AI
Alignment & Safety
Evaluation & Benchmarks
Mechanistic Interpretability
Reasoning
Research Paper
Latest papersRecent research connected to this question, newest first.Not All LLM Reasoning is Visible in the Chain-of-ThoughtThe evidence covers 13 frontier language models across three synthetic reasoning tasks, including filler-token effects and a hidden modular-arithmetic constraint in Claude Opus 4.5. It also examines reinforcement learning and supervised fine-tuning for Qwen3-235B; the findings establish invisible computation in these tested settings rather than across all models or tasks.research paper · Sep 3, 2026
Related questions
How reliably can chain-of-thought text reveal which reasoning steps causally drive correct answers?Can chain-of-thought monitoring detect preferences received through tools or inferred from raw artifacts?How can we test whether language models genuinely execute multi-step graph logic when static benchmarks become contaminated?How can we distinguish decodable logical validity from reasoning that actually drives a language model’s answers?