Get Started
Home
Topics
Search
Library
Research questionHow can static image encoders learn coherent object instances from unlabeled video?Many image-pretraining methods capture category semantics without preserving the identity and pixel-level coherence of individual instances. This gap limits transfer to monocular depth, 3D detection, occupancy prediction, and end-to-end planning.
AI
Computer Vision
Image & Video Processing
Machine Learning
Research Paper
Latest papersRecent research connected to this question, newest first.Object Concepts Emerge from MotionThe source studies motion-derived pseudo-instance supervision for single-image encoders using raw driving and web videos, without human annotations or camera calibration. Evidence covers transfer to monocular depth estimation, 3D object detection, 3D occupancy prediction, and end-to-end planning against supervised and self-supervised pretraining baselines.research paper · Sep 3, 2026
Related questions
How can RGB-D visual pretraining improve 3D awareness without sacrificing semantic transfer?How can video models recognize unseen actions from only a few labeled examples?How can surgical vision models learn scene geometry during pretraining while keeping inference RGB-only?How should self-supervised visual learning combine objectives to prevent collapse while preserving semantic and spatial information?