Get Started
Home
Topics
Search
Library
Research questionHow can robot perception encode action-relevant scene dynamics to improve manipulation generalization?Robot policies may recognize what is present without representing how objects and relevant regions change during manipulation. This can make visual control brittle when tasks or environments differ from those seen during learning.
AI
Computer Vision
Image & Video Processing
Machine Learning
Multimodal Models
Robotics
Latest papersRecent research connected to this question, newest first.DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided RepresentationThe source concerns an image-only visual encoder trained with image-language-3D-flow triplets from heterogeneous human and robot videos. Evidence covers downstream policies, including VLAs, in simulation and real-world settings, with reported gains in out-of-distribution scenarios.research paper · May 28, 2026
Related questions
How can robotic vision-language-action models generalize across backbones without losing hierarchical manipulation structure?How can pretrained vision-language-action models reliably perform contact-rich manipulation when goals, scenes, and contacts change?How can robot world-action models use 3D geometry to predict actions beyond RGB observations?How can robots safely learn dynamic manipulation skills online despite sim-to-real mismatch?