Get Started
Home
Topics
Search
Library
Research questionHow can surgical vision models learn scene geometry during pretraining while keeping inference RGB-only?Surgical datasets are often data-scarce, and RGB-only self-supervision can miss scene geometry that supports downstream understanding. Geometric training signals must therefore improve the representation without making inference depend on them.
AI
Computer Vision
Health
Image & Video Processing
Machine Learning
Multimodal Models
Research Paper
Latest papersRecent research connected to this question, newest first.DART: Depth-as-Target Pretraining for Surgical Vision Foundation ModelsThe evidence concerns DART, which extends DINOv2 with pseudo-labeled dense depth during pretraining and uses RGB only for fine-tuning and inference. Results cover eight surgical benchmarks for segmentation, depth estimation, and image-level recognition, with comparisons against natural-image and in-domain baselines.research paper · Sep 3, 2026
Related questions
How can RGB-D visual pretraining improve 3D awareness without sacrificing semantic transfer?How can surgical perception reconstruct 3D surfaces accurately from sparse endoscopic viewpoints?How can vision-language models infer 3D geometry and temporal continuity from 2D visual observations?How can RGB-D salient-object detection remain reliable when depth measurements are missing or corrupted?