Research questionHow can LLM sandbox security remain reliable when linguistic monitoring misrepresents internal computation?An LLM’s verbal self-reports and linguistically defined internal probes may not faithfully reflect its underlying activation-space computations. This makes it difficult to guarantee sandbox safety using language as the primary window into model behavior.