Get Started
Research questionHow can we test whether reference-based NLG metrics behave correctly under controlled response changes?Aggregate agreement can conceal whether a metric responds appropriately when a generated response undergoes a change that should or should not affect its score. Evaluators with similar overall performance may therefore make substantially different local judgments.
AI
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation MethodsThe evidence covers lexical, character-level, semantic, LLM-based, and hybrid reference-based evaluators. It examines correctness-preserving and correctness-altering response transformations, along with stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility.research paper · Sep 4, 2026
Related questions
How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?Can automatic metrics and LLM judges reliably reflect human judgments of multilingual summary quality?How can we audit NL-to-FOL benchmarks so annotation errors do not distort model evaluation?How can we reliably detect when an LLM response is unsupported by its reference documents?
Home
Topics
Search
Library