Research questionCan black-box LLM judges provide reproducible measurements on shared endpoints?The same request to the same model name may produce different rankings across repeated or later calls on shared infrastructure. This instability can make filtering, scoring, and pass/fail decisions irreproducible even when execution records are complete.