Get Started
Research questionHow does sparse mixture-of-experts routing trade off approximation, learning error, and computation under misspecification and evolving experts?Input-dependent gating lets an MoE specialize predictors across regions of the input space while limiting active computation. When experts evolve or fail to match the data-generating process, the effects of routing, local specialization, and shared predictive structure on overall risk remain unclear.
AI
Machine Learning
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.Towards a Statistical Understanding of Mixture-of-ExpertsThe source provides a theoretical framework with oracle risk bounds for dense and sparse routing with evolving experts. It separates approximation, expert-learning, and router-estimation errors, relates routing to local expert advantage, and analyzes how shared experts can capture common structure while routed experts model residual variation.research paper · Sep 3, 2026
Related questions
How should experts be pruned in over-dispersed MoE routing when router importance and perplexity mislead?How can inference prune activated MoE experts without confounding compute savings with output rescaling?When can factual recall in sparse MoE language models be attributed to one expert rather than an expert set?How can we tell whether cross-layer routing predictability reflects shared routing dynamics or generic hidden-state smoothness?
Home
Topics
Search
Library