Get Started
Research questionHow should LLM pre-training allocate a fixed token budget between repetition and auxiliary views when prior knowledge is incomplete?The same knowledge can appear repeatedly or through paraphrased reformulations, but it is unclear how these choices affect what language models learn. The trade-off may also vary with batch size and the type of knowledge being acquired.
LLM Pretraining & Post-training
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary ViewsApplies to controlled experiments on LLM pre-training examining repetition, paraphrased auxiliary views, fixed token budgets, teacher-model strength, prior knowledge gaps, and layer-wise or compression-related effects. The evidence concerns knowledge acquisition during pre-training rather than downstream deployment behavior.research paper · Sep 3, 2026
Related questions
How can we measure and control an LLM’s reliance on token-frequency priors when context is sparse?How should low-resource LLM fine-tuning use task-level language priors with ambiguous or incomplete data?How can instruction-tuned LLMs learn corpus-specific knowledge without exhaustive synthetic QA or instruction fine-tuning?How do pause tokens affect reasoning adaptation while preserving previously learned capabilities during LLM fine-tuning?
Home
Topics
Search
Library