Get Started
Home
Topics
Search
Library
Research questionHow much model capacity do LIBERO manipulation policies need to generalize beyond memorized task mappings?High LIBERO scores may reflect task-specific associations rather than genuine instruction following. The capacity needed for these tasks, and the extent to which performance survives task remapping or perturbations, is unclear.
AI
Evaluation & Benchmarks
Machine Learning
Multimodal Models
Robotics
Small / On-device Models
Latest papersRecent research connected to this question, newest first.MINERVA: How Small Can a Manipulation Policy Be and Still Solve LIBERO?The evidence covers compact vision-action policies evaluated on the four standard LIBERO suites, LIBERO-90, and LIBERO-Plus perturbations. It includes parameter scaling, a task-ID permutation probe, and CPU inference timing, but does not establish performance beyond the reported benchmark tasks.research paper · Sep 3, 2026
Related questions
How can robots learn robust, generalist bimanual household manipulation from limited human demonstrations?How can robot manipulation policies be evaluated for execution quality without costly, unstable repeated hardware trials?How can robotic vision-language-action models generalize across backbones without losing hierarchical manipulation structure?How can dexterous robot hands generalize a manipulation skill from one demonstration to varied objects despite sim-to-real gaps?