Get Started
Research questionCan automatic metrics and LLM judges reliably reflect human judgments of multilingual summary quality?Automatic evaluation makes it practical to compare summaries, but its scores may not capture how people judge summary quality. This makes it difficult to know which evaluators can be trusted across languages and criteria.
AI
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond EnglishThe evidence comes from BASSE, a multilingual meta-evaluation dataset containing human ratings of 2,040 abstractive summaries produced manually or by five LLMs with four prompts. Ratings cover coherence, consistency, fluency, relevance, and 5W1H; the study benchmarks automatic metrics and LLM judges, finding higher human-judgment correlation for proprietary judges, followed by criteria-specific metrics, while open-source judges perform poorly.research paper · Sep 2, 2026
Related questions
How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?How can we evaluate LLM reasoning quality beyond final-answer accuracy across deployment contexts?How can LLM judges produce reliable, unbiased scores for subjective, interdependent multi-step creativity tasks?Can black-box LLM judges provide reproducible measurements on shared endpoints?
Home
Topics
Search
Library