Get Started
Research questionHow can reward models distinguish fine-grained translation quality across candidate groups during GRPO post-training?Independent quality scores evaluate translations in isolation, making subtle differences between candidates difficult to resolve. This can weaken the relative rewards used to improve open-ended machine translation.
LLM Pretraining & Post-training
Natural Language Processing
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.GRRM: Group Relative Reward Modeling for Machine TranslationApplies to open-ended machine translation systems using GRPO, where candidate groups must be ranked by relative quality. The source provides evidence from reward-ranking and translation-quality evaluations, along with reported reasoning outcomes, code, datasets, and model checkpoints.research paper · Sep 1, 2026
Related questions
How can we evaluate machine translation reliably and actionably as standard benchmarks saturate?How can GRPO avoid reinforcing guessed correct answers in bounded-answer and search tasks?How should machine translation research balance benchmark accuracy with stakeholders’ trust and quality needs?How can DeepResearch systems obtain scalable, query-specific reward signals for report quality?
Home
Topics
Search
Library