Get Started
Home
Topics
Search
Library
Research questionHow can evaluators distinguish missing knowledge from miscalibrated outputs in language models?A model may encode a correct judgment while an output threshold produces the wrong answer. Observing only the final response therefore cannot reliably distinguish missing knowledge from a faulty readout.
AI
Evaluation & Benchmarks
Machine Learning
Mechanistic Interpretability
Natural Language Processing
Reasoning
Latest papersRecent research connected to this question, newest first.When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language ModelsThe source studies logical-verification judgments using hidden-state probes, output-logit margins, forced-choice and free-form responses, unseen structures, matched foils, and a surface-cue-controlled task. Because hidden-state and logit analyses require model instrumentation, its black-box analysis also shows that surface features can mimic probing evidence; conclusions are therefore limited when only ordinary outputs are available.research paper · Sep 4, 2026
Related questions
How can we detect internally incoherent language-model forecasts before relying on them for consequential decisions?How can safety evaluations measure language-model behavior without triggering evaluation-aware changes in decisions?How can we distinguish decodable logical validity from reasoning that actually drives a language model’s answers?How can evaluations separate model capability from execution-harness capability?