Research questionHow can decode requests be routed in PD-disaggregated MoE serving when expert sets cause unequal latency?Different decode batches can activate different sets of experts, changing the amount of expert weight loading even when workers have similar request loads. This makes conventional load balancing insufficient for keeping decode latency consistent.