Get Started
Research questionHow can we tell whether cross-layer routing predictability reflects shared routing dynamics or generic hidden-state smoothness?Each sparse MoE layer has its own router, yet later expert choices can often be predicted from earlier routing signals. This predictability may arise from generic smoothness in hidden representations rather than information specific to routing decisions.
AI
Machine Learning
Mechanistic Interpretability
Latest papersRecent research connected to this question, newest first.Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-ExpertsThe evidence concerns sparse MoE language models with separately parameterized routers across depth. It compares router-control states with matched-rank residual representations and tests whether predicted states preserve local routing behavior and language-model loss in OLMoE and Phi; these results do not establish the distinction for all MoE architectures.research paper · Sep 2, 2026
Related questions
How does sparse mixture-of-experts routing trade off approximation, learning error, and computation under misspecification and evolving experts?How should experts be pruned in over-dispersed MoE routing when router importance and perplexity mislead?When can factual recall in sparse MoE language models be attributed to one expert rather than an expert set?Which routing and mixing behaviors in trained four-stream residual pathways materially affect model performance?
Home
Topics
Search
Library