Research questionWhen should self-supervised pretraining pool dependent augmentations rather than partition data into independent subsets?Augmentations generated from the same unlabeled example are statistically dependent, complicating the choice between reusing them and restricting training to independent subsets. The trade-off depends on how those correlations affect estimation variance and the learned invariant subspace.