Research questionCan intermediate LLM activations guide faster jailbreak search without weakening attack effectiveness?Refusal behavior may be represented in transformer activations before the model produces its output. The difficulty is using that signal to reduce the cost of prompt search without losing the effectiveness of the resulting attacks.