Get Started
Home
Topics
Search
Library
Research questionHow can attention-head contributions be measured in prompt-injection classifiers across circuit and output scales?Many attention heads can jointly shape a classifier’s logits, while global output behavior can obscure which local circuit components drove the decision. The challenge is to connect fine-grained head behavior with the model’s final classification.
AI
Alignment & Safety
Machine Learning
Mechanistic Interpretability
Natural Language Processing
Research Paper
Latest papersRecent research connected to this question, newest first.Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference timeThe evidence concerns a DeBERTa model specialized for prompt-injection detection and examines attention-head behavior in correct and erroneous predictions at head, layer, and network scales. The supplied evidence does not establish generality beyond this classifier.research paper · Sep 4, 2026
Related questions
How can attention heads be pruned in text-to-image diffusion transformers without losing prompt-specific object identity?Do multilingual attention heads that retrieve context also control transitions into the target language during reasoning?How can we predict when added attention will help GI endoscopy classifiers across domain gaps?How many attention heads are necessary for Boolean computation in one-layer attention-only models, even with unlimited dimension and precision?