Research questionHow can we compare language models’ conditional behavior and predict the effects of prompt changes?Model-level summaries can obscure how a language model’s response distribution varies with the prompt. This makes it difficult to connect model differences and prompt changes to downstream task performance.