Get Started
Research questionHow can reliable foundation-model scaling laws be constructed without training every configuration?Scaling-law construction typically requires an expensive grid of training configurations, even though most configurations do not determine the best observed loss at each compute scale. The difficulty is identifying which runs provide enough information for reliable fits under a limited compute budget.
LLM Pretraining & Post-training
Machine Learning
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.Amortizing Scaling Law Construction CostsThe source studies a Bayesian-optimization framework that expands the compute budget progressively and uses surrogate-fantasized evaluations to recover the broader experimental grid. It reports scaling-law fits close to those from dense grids, with computational savings of up to 10–100×.research paper · Sep 4, 2026
Related questions
How can LLM pretraining avoid sudden gradient explosions when scaling to larger models?Can tabular foundation models learn transferable physical laws with units and noiseless mechanisms, not just interpolate data?How can validation loss track language-model capabilities across changing training distributions without benchmark evaluation?How can large language models cut training and inference costs without materially harming accuracy?
Home
Topics
Search
Library