Research questionHow should repeated-query audits determine whether LLM brand recommendations are reliable amid sampling and other sources of variation?Identical prompts can produce different brand recommendations, making apparent preferences difficult to distinguish from stochastic generation, prompt phrasing, run-to-run, or model-version effects. Auditors need evidence for reliability judgments without conflating these sources of variation.