Get Started
Research questionHow should optimizer selection and hyperparameter tuning account for longer training horizons?An optimizer that is effective at one training duration may rank differently or require different settings at another. Longer training also changes how memory, decay, and learning-rate schedules influence efficiency and final quality.
AI
Evaluation & Benchmarks
Machine Learning
Research Paper
Latest papersRecent research connected to this question, newest first.Optimizer Memory Schedules for Outscaling the Overtraining AxisThe source compares several optimizer families across multiple model sizes and overtraining horizons while retuning learning rates and related settings. Its conclusions are based on the reported model scales, horizons, and synthetic-theory analysis, and may not extend beyond them.research paper · Sep 4, 2026
Related questions
How can functional bilevel optimization adapt online as learning objectives change over time?How can policy optimization for long-horizon LLM agents preserve useful transitions across updates when rollout groups are small?How can LoRA fine-tuning optimize fixed-rank updates while respecting the induced weight-matrix geometry?How can neural networks retain capacity while fitting within fixed parameter and memory budgets?
Home
Topics
Search
Library