Get Started
Home
Topics
Search
Library
Research questionHow can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?LLM judges can agree because they recognize quality, but they can also share biases that diverge from human evaluations. Agreement among judges alone cannot distinguish these explanations.
AI
Alignment & Safety
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human AlignmentThe evidence covers 42 judges evaluated on two community-built Indic benchmarks spanning four domains and eight languages. Comparisons use the same two-rater human reference, with analyses distinguishing subjective rubrics from one rubric with verifiable answers; the reported evidence includes judge–judge, judge–human, and human–human agreement as well as score-space geometry.research paper · Sep 2, 2026
Related questions
How can users judge whether an individual LLM recommendation merits reliance without objective ground truth?How can LLM graders assign accurate marks while grounding each judgment in rubrics and student-answer evidence?Can automatic metrics and LLM judges reliably reflect human judgments of multilingual summary quality?How can we evaluate LLM reasoning quality beyond final-answer accuracy across deployment contexts?