Research questionHow should math-focused retrievers be evaluated when generic benchmarks miss fine-grained relevance?Mathematical retrieval often depends on symbolic structure and subtle solution-level relevance that broad text-retrieval benchmarks may not capture. Strong performance on a generic benchmark therefore may not indicate useful retrieval for math-focused downstream systems.