Get Started
Home
Topics
Search
Library
Research questionHow can libraries compare bibliographic methods across tasks under limited computing budgets?Libraries and archives must evaluate methods across varied bibliographic tasks, but existing benchmarks do not consistently represent this work or the resources methods require. Performance can also differ substantially across bibliographic facets and tasks.
AI
Evaluation & Benchmarks
Information Retrieval
Machine Learning
Natural Language Processing
Latest papersRecent research connected to this question, newest first.SHELF: A Synthetic Harness for Multi-Task Bibliographic BenchmarkingSHELF provides synthetic benchmark data for classification, clustering, retrieval, pair classification, and instruction retrieval, comparing sparse methods, BM25, encoders, and selected zero-shot decoders. Its evidence includes task scores, a subject-classification timing comparison, and comparisons with other datasets. The synthetic scores do not estimate production catalogue accuracy, although method rankings transfer more reliably than absolute scores.research paper · Sep 2, 2026
Related questions
How can LLM agents retrieve relevant skills from large, noisy libraries under context and latency constraints?How can intelligent systems be compared under deployment constraints when representational economy, prediction, and resource use trade off?How can scattered findings about rapidly changing AI models become searchable and comparable?How can production RAG teams maintain reliable comparisons as new retrieval candidates arrive without rejudging overlapping documents?