Get Started
Home
Topics
Search
Library
Research questionHow can we audit NL-to-FOL benchmarks so annotation errors do not distort model evaluation?NL-to-FOL benchmarks can contain incorrect logical formalizations, ambiguous natural-language statements, and incorrect inference labels. These defects can make benchmark scores reflect reference-data errors rather than model capabilities.
AI
Evaluation & Benchmarks
Natural Language Processing
Reasoning
Latest papersRecent research connected to this question, newest first.Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human RelabelingThe evidence covers systematic human inspection of the FOLIO validation split and a subset of MALLS test instances, corrected ground truths, and an LLM-assisted framework for prioritizing manual review. The reported results include 11–23 percentage-point accuracy changes for three tested LLMs and 90% dataset accuracy after reviewing fewer than 20% of instances, compared with more than 76% for unguided review; findings are limited to the examined datasets, splits, subsets, and models.research paper · Sep 3, 2026
Related questions
How should NLP papers report human annotators and quality controls for valid, reproducible results?How can we test whether reference-based NLG metrics behave correctly under controlled response changes?How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?How can we evaluate machine translation reliably and actionably as standard benchmarks saturate?