Get Started
Research questionHow should NLP papers report human annotators and quality controls for valid, reproducible results?NLP studies rely on human annotations, but papers often omit who produced them and how annotation quality was controlled. These omissions make it difficult to judge the validity and reproduce the resulting evidence.
AI
Evaluation & Benchmarks
Natural Language Processing
Research Paper
Latest papersRecent research connected to this question, newest first.Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025The evidence covers 2,667 annotation tasks in 1,603 ACL-venue papers published from 2018 to 2025. It uses a taxonomy of reporting practices and an LLM-assisted extraction pipeline validated against human-adjudicated annotations from 41 papers and 72 tasks.research paper · Sep 2, 2026
Related questions
How can we audit NL-to-FOL benchmarks so annotation errors do not distort model evaluation?How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?How can we test whether reference-based NLG metrics behave correctly under controlled response changes?How can we measure whether computer-vision models are understandable to independent human evaluators?
Home
Topics
Search
Library