Get Started
Home
Topics
Search
Library
Research questionHow can document retrieval find the correct record when large collections share nearly identical visual templates?Enterprise repositories may contain many records with nearly identical layouts but different contents. This visual similarity can make embeddings insufficiently discriminative, causing retrieval to select the wrong document.
AI
Computer Vision
Evaluation & Benchmarks
Finance
Information Retrieval
Multimodal Models
Retrieval-Augmented Generation
Technology
Latest papersRecent research connected to this question, newest first.Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual HomogeneityThe evidence concerns Invoice Haystack, a benchmark of 1,500 anonymized invoice images with 200 discriminative question-answer pairs, and reports results for hybrid text-visual retrieval with VLM-based verification. Findings are demonstrated on invoice collections and comparisons with DocHaystack and InfoHaystack; they do not establish performance across all enterprise document types or deployment settings.research paper · Sep 3, 2026
Related questions
How can visual-document RAG adapt page retrieval to each query without hurting answer accuracy?How can cross-modal generative retrieval avoid hallucinating visual details missing from concise text queries?How can multimodal retrieval distinguish correct attribute–object bindings when images share the same concepts?How can image-text retrieval focus on caption-described attributes while ignoring unmentioned visual information?