Research questionHow can LLM hate-speech moderation resist annotator-style rebuttals that reverse correct judgments across repeated interactions?In human–AI moderation workflows, annotator-style rebuttals can overturn an initially correct LLM judgment by normalizing hateful content or labeling harmless content hateful. Repeated exchanges may amplify this instability, and the two reversal directions need not affect models equally.