Get Started
Research questionHow can small-scale pretraining mixture experiments stay reliable when scarce high-quality data is repeated at target scale?When high-quality sources are small, increasing the training budget changes how often those examples recur. A mixture that looks optimal in a small experiment can therefore become suboptimal when extrapolated to the full training run.
AI
LLM Pretraining & Post-training
Machine Learning
Latest papersRecent research connected to this question, newest first.Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix ThemThe evidence covers two-source mixtures of limited high-quality data and web crawl, plus a three-source setting, at 757M and 1.17B parameter scales. Results include Wiki-Text evaluation and comparisons of token budgets for repetition-controlled versus uncontrolled experiments; broader model scales and mixture types are not established.research paper · Sep 3, 2026
Related questions
When and why can repeating a smaller dataset reduce training compute versus using more unique samples?How should LLM pre-training allocate a fixed token budget between repetition and auxiliary views when prior knowledge is incomplete?How can model distillation block hidden teacher-trait transfer through clean data without degrading the target task?When should self-supervised pretraining pool dependent augmentations rather than partition data into independent subsets?
Home
Topics
Search
Library