Get Started
Home
Topics
Search
Library
Research questionHow should repeated-query audits determine whether LLM brand recommendations are reliable amid sampling and other sources of variation?Identical prompts can produce different brand recommendations, making apparent preferences difficult to distinguish from stochastic generation, prompt phrasing, run-to-run, or model-version effects. Auditors need evidence for reliability judgments without conflating these sources of variation.
AI
Evaluation & Benchmarks
Machine Learning
Natural Language Processing
Research Paper
Statistical Machine Learning
Technology
Latest papersRecent research connected to this question, newest first.The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand RecommendationsThe study covers repeated-query audits of LLM brand recommendations under temperature-scaled nucleus sampling. It analyzes approximately 190,000 observations across more than 270 brands, six languages, and five to 40 iterations, using multiple metric families, variance decomposition, bootstrap and simulation analyses, and drift diagnostics on pinned snapshots. External validation reproduced reliability predictions in 37 of 39 cells, but fixed iteration tiers did not transfer, limiting the evidence for universal repeat-count rules.research paper · Sep 3, 2026
Related questions
How can we measure saturation of repeated LLM recommendations without fixed rosters masking new brands and sources?How can users judge whether an individual LLM recommendation merits reliance without objective ground truth?How can audits separate genuine LLM-judge preference effects from artifacts of double differences on bounded ratings?How can we evaluate LLM reasoning quality beyond final-answer accuracy across deployment contexts?