Research questionHow can LLM pretraining avoid sudden gradient explosions when scaling to larger models?Pretraining can abruptly fail when gradients explode, wasting the computation invested before the collapse. The failure is preceded by declining weight-matrix stable rank and increasing alignment between adjacent-layer Jacobians.