Get Started
Research questionHow can content moderation benchmarks reveal whether LLMs apply the intended criterion for each decision?A single moderation label may combine several criteria, so strong aggregate accuracy can conceal whether a model evaluates the relevant aspect of the content. A model may instead rely on its overall impression of harmfulness.
AI
Alignment & Safety
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content ModerationEvidence spans four moderation datasets and four LLMs, showing that models can perform strongly on aggregated benchmarks while failing when correct decisions depend on a specific content aspect.research paper · Sep 3, 2026
Related questions
How can we evaluate LLM reasoning quality beyond final-answer accuracy across deployment contexts?How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?How can LLM graders assign accurate marks while grounding each judgment in rubrics and student-answer evidence?Can black-box LLM judges provide reproducible measurements on shared endpoints?
Home
Topics
Search
Library