Get Started
Home
Topics
Search
Library
Research questionHow can robot world-action models use 3D geometry to predict actions beyond RGB observations?RGB-only robot world-action models do not directly represent scene geometry, which can make spatial relationships harder to use when predicting future states and actions. Adding geometric information must also preserve the visual and physical priors learned from large-scale video.
AI
Computer Vision
Diffusion Models
Evaluation & Benchmarks
Multimodal Models
Robotics
Latest papersRecent research connected to this question, newest first.Spatially Aware World Action Model via Geometric Latent DiffusionThis concerns world-action models built from pretrained video diffusion models for robotics. The source describes a model that jointly predicts RGB observations, depth, and actions with one diffusion backbone, reuses a frozen VAE tokenizer through nonlinear depth encoding, and evaluates robot control on RoboCasa, LIBERO-Plus, and a UR5 arm in real-world randomized environments.research paper · Sep 2, 2026
Related questions
How can robot perception encode action-relevant scene dynamics to improve manipulation generalization?How can vision-language models infer 3D geometry and temporal continuity from 2D visual observations?How can robotic vision-language-action models generalize across backbones without losing hierarchical manipulation structure?How can RGB-D visual pretraining improve 3D awareness without sacrificing semantic transfer?