Get Started
Home
Topics
Search
Library
Research questionHow can audits separate genuine LLM-judge preference effects from artifacts of double differences on bounded ratings?Double differences on bounded ratings can turn a common severity shift into an apparent interaction when candidate responses are censored unequally. The observed effect may therefore conflate differential preference with differential attenuation.
AI
Evaluation & Benchmarks
Machine Learning
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge AuditThe evidence concerns a pre-registered audit of a frozen pedagogy judge using 990 calls, testing whether a stated learner profile changed the judge’s preference for scaffolding. It derives the censoring mechanism and measures its contribution from the audit’s own ratings; the evidence is specific to this judge and task.research paper · Sep 2, 2026
Related questions
How should repeated-query audits determine whether LLM brand recommendations are reliable amid sampling and other sources of variation?How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?Can black-box LLM judges provide reproducible measurements on shared endpoints?How reliably do proxy LLM judges capture users’ perceived helpfulness and privacy in privacy-sensitive scenarios?