Get Started
Home
Topics
Search
Library
Research questionHow can RGB-D visual pretraining improve 3D awareness without sacrificing semantic transfer?RGB-only pretraining lacks the explicit geometric information provided by depth sensors, limiting 3D understanding in embodied systems. Incorporating depth must also retain features useful for semantic tasks.
AI
Computer Vision
Image & Video Processing
Machine Learning
Multimodal Models
Robotics
Latest papersRecent research connected to this question, newest first.DINOcular: Self-Supervised Visuospatial RepresentationsThe source presents a self-supervised RGB-D representation framework that fuses depth-derived geometric priors with a visual backbone through inter-patch and intra-patch fusion. Evidence covers 3D geometry benchmarks and RGB-D semantic segmentation tasks; it does not establish performance beyond those settings.research paper · Sep 3, 2026
Related questions
How can surgical vision models learn scene geometry during pretraining while keeping inference RGB-only?How can we predict local 3D occupancy from a single underwater RGB image despite unreliable visual geometry?How can robot world-action models use 3D geometry to predict actions beyond RGB observations?How can RGB-D salient-object detection remain reliable when depth measurements are missing or corrupted?