Get Started
Home
Topics
Search
Library
Research questionHow can LLM hate-speech moderation resist annotator-style rebuttals that reverse correct judgments across repeated interactions?In human–AI moderation workflows, annotator-style rebuttals can overturn an initially correct LLM judgment by normalizing hateful content or labeling harmless content hateful. Repeated exchanges may amplify this instability, and the two reversal directions need not affect models equally.
AI
Alignment & Safety
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based ModerationThe evidence concerns multiple LLMs evaluated on two hate-speech datasets using a rejudge protocol with direct contradiction, decision-boundary perturbations, and adversarial rationales. It reports stronger effects in multi-turn settings, model-specific asymmetries between whitewashing and smearing, and partial—not complete—mitigation from explicit reasoning prompts and defensive instructions.research paper · Sep 10, 2026
Related questions
Can user feedback reliably guide LLM revisions if LLM judges overlook the resulting improvements?Do source labels bias human and LLM judgments of logical fallacies differently?How can LLM prompts be automatically refined from recurring reasoning errors without laborious manual engineering?How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?