Get Started
Home
Topics
Search
Library
Research questionHow can leaderboards quantify uncertainty from sparse, unevenly sampled pairwise LLM judgments?Pairwise human judgments provide noisy, incomplete evidence about latent model abilities, and comparisons are not sampled uniformly. Consequently, leaderboard scores may obscure how reliable estimated differences or win probabilities are.
AI
Evaluation & Benchmarks
Machine Learning
Research Paper
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric EfficiencyThe analysis applies to LLM evaluation platforms using pairwise human judgments under Bradley–Terry–Luce-type models. The source studies semiparametric efficiency and a debiased estimator with asymptotic normality for these low-rank structures; it does not identify a specific platform or deployment setting.research paper · Sep 3, 2026Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise ComparisonsThe evidence concerns task-specific LLM rankings inferred from pairwise human-preference data, including theoretical guarantees and experiments on synthetic data and Chatbot Arena. It covers uncertainty for score contrasts, per-task ranks, and top-K membership, but does not establish behavior beyond these settings.research paper · Sep 2, 2026
Related questions
When can a random projection preserve JL distance bounds yet lose geometric variation needed for rankings and inference?When does an LLM’s verbal confidence reliably reflect its underlying uncertainty?How can LLMs distinguish ambiguous inputs from gaps in their knowledge when estimating uncertainty?How can LLMs produce reliable confidence estimates for deciding when to defer outputs to humans?