Research questionHow can surgical vision models learn scene geometry during pretraining while keeping inference RGB-only?Surgical datasets are often data-scarce, and RGB-only self-supervision can miss scene geometry that supports downstream understanding. Geometric training signals must therefore improve the representation without making inference depend on them.