Research questionHow does sparse mixture-of-experts routing trade off approximation, learning error, and computation under misspecification and evolving experts?Input-dependent gating lets an MoE specialize predictors across regions of the input space while limiting active computation. When experts evolve or fail to match the data-generating process, the effects of routing, local specialization, and shared predictive structure on overall risk remain unclear.