Research questionHow should language-model robustness be evaluated when text perturbations affect hidden states and attention differently from outputs?Typos, altered words, corrupted text, and disrupted token order can change a model’s internal computation without being fully reflected in its output behavior. Different perturbations may also produce distinct effects across hidden-state geometry and attention-head function.