Research questionHow can attention-head contributions be measured in prompt-injection classifiers across circuit and output scales?Many attention heads can jointly shape a classifier’s logits, while global output behavior can obscure which local circuit components drove the decision. The challenge is to connect fine-grained head behavior with the model’s final classification.