Get Started
Home
Topics
Search
Library
Research questionHow can vision-language reward models remain consistent when equivalent robot goals are paraphrased?A reward model may judge the same robot behavior differently solely because a goal description uses different wording. Such inconsistency can turn identical behavior into opposite progress or success assessments, undermining reward-guided learning.
AI
Computer Vision
Evaluation & Benchmarks
Machine Learning
Multimodal Models
Natural Language Processing
Reinforcement Learning
Robotics
Latest papersRecent research connected to this question, newest first.Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward ModelsThe evidence comes from ROBORMBENCH: 2,390 real-robot trajectories with ground-truth progress labels and 21,673 verified paraphrases covering lexical, syntactic, and action-goal rewrites. It examines proprietary and open-source VLMs, reports widespread instability that increases with rewrite divergence, and finds that trajectory-grounded reward models are more stable; scale and explicit reasoning do not reliably resolve the issue.research paper · Sep 4, 2026
Related questions
How can vision-language-action policies follow execution details beyond a robot task’s goal?How can we diagnose vision-language-action models’ failures on spatially ambiguous, long-horizon manipulation tasks?How can pretrained vision-language-action models reliably perform contact-rich manipulation when goals, scenes, and contacts change?How can closed-loop vision-language navigation learn effectively despite distribution shift and sparse micro-action rewards?