Get Started
Research questionCan black-box LLM judges provide reproducible measurements on shared endpoints?The same request to the same model name may produce different rankings across repeated or later calls on shared infrastructure. This instability can make filtering, scoring, and pass/fail decisions irreproducible even when execution records are complete.
AI
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared EndpointsThe evidence concerns black-box LLM judges accessed through shared endpoints, using repeated same-window and next-day replays across multiple providers, with a limited self-hosted comparison. It reports externally measured reproducibility and does not establish behavior beyond the tested infrastructure and configurations.research paper · Sep 3, 2026
Related questions
How reliably do proxy LLM judges capture users’ perceived helpfulness and privacy in privacy-sensitive scenarios?How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?How can LLMs generate reliable, adaptive tests that expose one another’s model-specific weaknesses?How should a single anchor be chosen to produce reliable rankings in LLM evaluation?
Home
Topics
Search
Library