Research questionWhen do tied or untied attention parameterizations enable weak recovery under stochastic training?High-dimensional attention models can exhibit uninformative optimization states, and different parameterizations can create distinct symmetry-breaking and timescale dynamics. The population-loss description may also differ substantially from the moment hierarchy governing online stochastic training.