Get Started
Home
Topics
Search
Library
Research questionDo source labels bias human and LLM judgments of logical fallacies differently?Labels about who produced a comment can influence judgments of its credibility and logical quality independently of the argument itself. It remains unclear whether humans and language models are affected by these source cues in the same way.
AI
Alignment & Safety
Evaluation & Benchmarks
Natural Language Processing
Reasoning
Latest papersRecent research connected to this question, newest first.Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMsEvidence comes from an online study of 505 participants evaluating comments with logical fallacies under human, AI, human-assisted, AI-assisted, and undisclosed source conditions. GPT-5.2, Gemini 2.5 Flash, and Claude Sonnet 4.5 were evaluated under the same conditions; the findings report stronger source-label effects among humans and comparatively stable, model-dependent responses among LLMs.research paper · Sep 4, 2026
Related questions
How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?How can human reviewers reliably detect LLM errors when verification reasoning is hard to retrieve?Can user feedback reliably guide LLM revisions if LLM judges overlook the resulting improvements?How can LLMs produce reliable confidence estimates for deciding when to defer outputs to humans?