Research questionHow can we audit NL-to-FOL benchmarks so annotation errors do not distort model evaluation?NL-to-FOL benchmarks can contain incorrect logical formalizations, ambiguous natural-language statements, and incorrect inference labels. These defects can make benchmark scores reflect reference-data errors rather than model capabilities.