Research questionCan pre-generalization interventions reveal when training constrains which equally fitting neural-network solutions will later generalize?Overparameterized networks can fit the same training data while differing substantially on unseen examples. During grokking, generalization emerges after a plateau, making it difficult to determine when training has begun constraining the eventual solution.