Get Started
Home
Topics
Search
Library
Research questionHow can molecular property benchmarks distinguish genuine prediction from verbatim retrieval of published values?A benchmark score can appear accurate when a model reproduces a molecular property value encountered during pretraining rather than inferring it. Retrieval varies across datasets and reasoning settings, making model comparisons difficult.
AI
Evaluation & Benchmarks
Information Retrieval
Machine Learning
Latest papersRecent research connected to this question, newest first.Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language ModelsThe evidence comes from an audit of 22 frontier language models across 12 regression benchmarks, testing verbatim retrieval at two reasoning levels. It also examines retrieval involving transformed SMILES strings and original labels, but does not establish a universal retrieval-free evaluation protocol.research paper · Sep 4, 2026
Related questions
How can molecular property predictors retain substructure and graph-distance information in compact representations without external pretraining?When do self-supervised molecular graph representations improve property prediction beyond Morgan fingerprints?Can prompt phrasing reliably improve LLM-derived chemical features for drug-toxicity prediction?How can we evaluate LLM reasoning quality beyond final-answer accuracy across deployment contexts?