Research questionHow can small-scale pretraining mixture experiments stay reliable when scarce high-quality data is repeated at target scale?When high-quality sources are small, increasing the training budget changes how often those examples recur. A mixture that looks optimal in a small experiment can therefore become suboptimal when extrapolated to the full training run.