Get Started
Home
Topics
Search
Library
Research questionWhen do tied or untied attention parameterizations enable weak recovery under stochastic training?High-dimensional attention models can exhibit uninformative optimization states, and different parameterizations can create distinct symmetry-breaking and timescale dynamics. The population-loss description may also differ substantially from the moment hierarchy governing online stochastic training.
AI
Machine Learning
Research Paper
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.High-Dimensional Learning Dynamics of Attention-Indexed ModelsThe source studies attention-indexed models representing multi-layer and multi-head attention. It analyzes direct matrices S, tied factorizations S=WWᵀ, and untied factorizations S=UVᵀ in a high-dimensional limit, characterizing population loss through trace order parameters and online SGD through matrix moments. Its recovery results concern weak recovery on the Θ(d² log d) sample scale; for untied attention, recovery depends on whether fast dynamics select a symmetry-breaking state.research paper · Sep 3, 2026
Related questions
How can language-model attention remain reliable beyond its training context?How can LoRA initialization preserve full-rank training gradients despite its low-rank bottleneck?Which attention mechanisms reliably improve DeepONet accuracy across data-driven and physics-informed PDE solving?How can offline preference optimization identify which chosen–rejected pairs merit gradients without destabilizing reasoning-model training?