Research questionHow can static image encoders learn coherent object instances from unlabeled video?Many image-pretraining methods capture category semantics without preserving the identity and pixel-level coherence of individual instances. This gap limits transfer to monocular depth, 3D detection, occupancy prediction, and end-to-end planning.