Research questionHow should LLM pre-training allocate a fixed token budget between repetition and auxiliary views when prior knowledge is incomplete?The same knowledge can appear repeatedly or through paraphrased reformulations, but it is unclear how these choices affect what language models learn. The trade-off may also vary with batch size and the type of knowledge being acquired.