Get Started
Home
Topics
Search
Library
Research questionHow can pretrained vision-language-action models reliably perform contact-rich manipulation when goals, scenes, and contacts change?Pretrained vision-language-action models can perform local manipulation skills but often fail when the requested goal or scene differs from their training trajectories. Analytic planners provide compositional control yet remain unreliable for irregular grasps, constrained placements, and articulated-object interactions.
AI
AI Agents
Multimodal Models
Reasoning
Robotics
Latest papersRecent research connected to this question, newest first.Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided AgentsThe source studies a frozen VLA exposed as a retryable contact-rich primitive alongside a fixed library of analytic primitives and memory from execution traces. Evidence covers perturbed tabletop, household kitchen, and clean-to-randomized bimanual manipulation tasks evaluated on LIBERO-Pro, RoboCasa365, and RoboTwin C2R; it does not establish broader deployment performance.research paper · Sep 2, 2026
Related questions
How can we diagnose vision-language-action models’ failures on spatially ambiguous, long-horizon manipulation tasks?How can robotic manipulation policies combine vision, language, and touch for reliable contact-rich tasks under occlusion?How can robotic vision-language-action models generalize across backbones without losing hierarchical manipulation structure?How can vision-language-action systems reliably execute long-horizon manipulation while tracking state and conditional dependencies?