Research questionHow can inference prune activated MoE experts without confounding compute savings with output rescaling?Reducing the number of selected experts changes both which computations are performed and how router probabilities are normalized. Quality loss may therefore result from altered output strength rather than expert removal alone.