Get Started
Home
Topics
Search
Library
Research questionHow can we evaluate machine translation reliably and actionably as standard benchmarks saturate?Standard translation benchmarks increasingly distinguish poorly among stronger systems, while automatic metrics can be unreliable or reward-hacked. Human evaluation can identify problems but may lack reproducibility, objectivity, and scalability.
AI
Evaluation & Benchmarks
Machine Learning
Multimodal Models
Natural Language Processing
Research Paper
Technology
Latest papersRecent research connected to this question, newest first.Beyond BLEU: A Case for Redefining Sign Language Translation BenchmarksThe evidence covers six models evaluated on Phoenix-2014T and CSL-Daily. It compares BLEU-4 with an open-weight-LLM question-answering protocol measuring salient content preservation, including its relationship to human rankings, paraphrase invariance, train-test overlap, and differences between gloss-free and gloss-supervised systems.research paper · Sep 3, 2026Last Translation BenchmarkThe source describes a live benchmark of human-authored, peer-reviewed examples involving text, images, audio, and video. Each example includes handcrafted verification rules for specific failure cases; the evidence establishes the benchmark and evaluation design, not broader performance claims.research paper · Sep 3, 2026Is Human Annotation Necessary? Iterative MBR Distillation for Error Span Detection in Machine TranslationThe source studies iterative MBR distillation using an off-the-shelf LLM to generate pseudo-labels for training without human annotations. It evaluates the approach on WMT Metrics Shared Task datasets at system, span, and sentence levels.research paper · Sep 1, 2026
Related questions
How should machine translation research balance benchmark accuracy with stakeholders’ trust and quality needs?How can we audit NL-to-FOL benchmarks so annotation errors do not distort model evaluation?How can reward models distinguish fine-grained translation quality across candidate groups during GRPO post-training?How can we test whether reference-based NLG metrics behave correctly under controlled response changes?