Get Started
Research questionWhen should self-supervised pretraining pool dependent augmentations rather than partition data into independent subsets?Augmentations generated from the same unlabeled example are statistically dependent, complicating the choice between reusing them and restricting training to independent subsets. The trade-off depends on how those correlations affect estimation variance and the learned invariant subspace.
Machine Learning
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.An Analysis of Self-supervised Pre-training with Dependent SamplesThe analysis compares pooled augmentations with partitioning data into independent subsets when estimating an invariant subspace. It reports bounds showing pooling is never worse in the stated setting and can yield faster rates for masking or noise-injection augmentations with a shallow neural network; broader architectures and augmentation schemes are not established by the source.research paper · Sep 4, 2026
Related questions
In multivariate forecasting, how can augmentation vary samples without breaking input–future temporal coherence?How can small-scale pretraining mixture experiments stay reliable when scarce high-quality data is repeated at target scale?When augmentation quantity is fixed, how does synthetic-example placement in representation space affect imbalanced pragmatic-function classification?How can synthetic image augmentation reliably improve computer vision when labeled real data are scarce?
Home
Topics
Search
Library