Get Started
Home
Topics
Search
Library
Research questionHow can vision-language-action policies act reliably when irrelevant sensors are corrupted or only one informative sensor remains?Limited, homogeneous robot demonstrations can make a VLA policy depend on cross-sensor correlations rather than the sensor carrying task-relevant information. Consequently, corrupting an irrelevant modality can disrupt actions, while the policy may fail when only one informative modality remains.
AI
Evaluation & Benchmarks
Machine Learning
Multimodal Models
Robotics
Latest papersRecent research connected to this question, newest first.Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA PoliciesThe problem is studied with BEHAVIOR-1K simulation diagnostics and 47 rollout-based skills, plus two physical setups: a bi-manual system with two Kinova arms and three RGB cameras, and a single-arm MELFA ASSISTA system with vision and GelSight tactile sensing. Evidence covers full modalities, uninformative-sensor corruption, single-sensor fallback, and physical-object distractors; it is limited to these benchmarks and robot embodiments.research paper · Sep 2, 2026
Related questions
How can robotic manipulation policies combine vision, language, and touch for reliable contact-rich tasks under occlusion?How can dual-arm vision-language-action policies avoid self-collisions with grasped objects?How can we diagnose vision-language-action models’ failures on spatially ambiguous, long-horizon manipulation tasks?How can pretrained vision-language-action models reliably perform contact-rich manipulation when goals, scenes, and contacts change?