Get Started
Research questionHow can production RAG teams maintain reliable comparisons as new retrieval candidates arrive without rejudging overlapping documents?Comparing retrieval systems requires relevance judgments over their retrieved documents, but candidate systems often return overlapping results. Reassessing those documents whenever a new candidate arrives makes ongoing model selection expensive and difficult to repeat.
Evaluation & Benchmarks
Information Retrieval
Retrieval-Augmented Generation
Latest papersRecent research connected to this question, newest first.Incremental Pooled LLM Evaluation for Cost-Effective Retrieval Model SelectionThe source studies pooled LLM relevance judgments, expanding the judged pool only with documents contributed by newly introduced retrieval systems. Evidence covers four retrieval benchmarks with 11 dense, sparse, and hybrid systems, plus 62 retrieval configurations in a financial news question-answering system; reported cost and ranking results are specific to these settings.research paper · Sep 2, 2026
Related questions
How can long-context RAG preserve global document structure during retrieval?How can visual-document RAG adapt page retrieval to each query without hurting answer accuracy?How can production recommender systems continually optimize retrieval, ranking, and serving as users and content change?How can structured RAG synthesize evidence scattered across many documents while limiting answer-generation cost?
Home
Topics
Search
Library