Research questionHow can leaderboards quantify uncertainty from sparse, unevenly sampled pairwise LLM judgments?Pairwise human judgments provide noisy, incomplete evidence about latent model abilities, and comparisons are not sampled uniformly. Consequently, leaderboard scores may obscure how reliable estimated differences or win probabilities are.