Research questionHow can production RAG teams maintain reliable comparisons as new retrieval candidates arrive without rejudging overlapping documents?Comparing retrieval systems requires relevance judgments over their retrieved documents, but candidate systems often return overlapping results. Reassessing those documents whenever a new candidate arrives makes ongoing model selection expensive and difficult to repeat.