Get Started
Research questionHow should momentum and batch size be tuned to preserve stability and data efficiency in one-pass training?In one-pass large-batch training, increasing batch size changes the learning-rate range that keeps risk stable and can alter data efficiency. Polyak and Nesterov momentum can produce different risk trajectories as batch size grows.
Machine Learning
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiencyThe analysis focuses on Polyak and Nesterov momentum in a tractable power-law kernel regression setting. It characterizes critical learning rates, risk dynamics, and batch-size regimes under a fixed data budget, with numerical experiments validating the reported stability boundaries and scaling behavior.research paper · Sep 2, 2026
Related questions
How should optimizer selection and hyperparameter tuning account for longer training horizons?How can small-scale pretraining mixture experiments stay reliable when scarce high-quality data is repeated at target scale?How can deep-network momentum adapt forgetting to unevenly sampled input directions?How can neural networks retain capacity while fitting within fixed parameter and memory budgets?
Home
Topics
Search
Library