Research questionHow can content moderation benchmarks reveal whether LLMs apply the intended criterion for each decision?A single moderation label may combine several criteria, so strong aggregate accuracy can conceal whether a model evaluates the relevant aspect of the content. A model may instead rely on its overall impression of harmfulness.