Research questionHow can evaluators distinguish missing knowledge from miscalibrated outputs in language models?A model may encode a correct judgment while an output threshold produces the wrong answer. Observing only the final response therefore cannot reliably distinguish missing knowledge from a faulty readout.