Research questionHow can safety evaluations measure language-model behavior without triggering evaluation-aware changes in decisions?Language models may respond differently when told that their alignment is being tested. This can alter both their overall decisions and the information they use to make them, making standard evaluations harder to interpret.