SlimWise runs MoE prefill with the full expert pool and decode with a pruned pool over a shared KV cache, lifting decode throughput by up to 1.81× at 50% expert pruning with minimal accuracy loss.
You’re serving a sparse Mixture of Experts model like Qwen3.6-35B-A3B. Each token only activates 8 of 256 experts, so you’d expect decoding to be cheap. In practice, when you batch dozens of user requests together, the union of experts chosen across the batch covers almost the whole pool. At batch size 64, the paper measures that each decode step touches roughly 75% of the 60 GiB expert pool. Decoding becomes memory-bound: you’re shuttling expert weights from GPU HBM every step.
The standard fix is expert pruning: rank experts by importance on a calibration set, keep the top ones, discard the rest. But conventional pruning applies the same reduced pool to both inference phases. Prefill (processing the prompt) is already compute-bound because many tokens share each loaded expert, so shrinking the pool costs accuracy and buys little speed. Pruning criteria like REAP can also warp generation behavior: on MATH-500, median output length jumps from 2.4k to 6.6k tokens even when benchmark accuracy looks fine, and on HumanEval+ responses can collapse to 0.3k tokens.
The core observation: prefill and decode have opposite bottlenecks, so they should get different expert pools. SlimWise runs prefill with every expert active, then hands the resulting KV cache directly to a pruned decoder. Because pruning doesn’t change the attention architecture or cache layout, no conversion is needed. The pruned decoder reads full-model keys and values as if it had produced them itself.
Under Prefill-Decode disaggregation, prefill and decode already live on separate instances, so SlimWise just loads a smaller checkpoint on the decode side. Freed memory becomes extra KV cache, which supports larger batches. For co-located serving, SlimWise expresses pruning as phase-aware router masking: during decode, the router’s logits for removed experts are set to a large negative value before top-k selection, so the engine keeps one copy of the full model but behaves as pruned when generating.
The training-free handoff recovers much of the lost accuracy but doesn’t fix the runaway-generation problem on math tasks. SlimWise adds a cheap distillation stage: train the pruned decoder to continue from teacher-generated KV caches. For each training sequence, sample a split point s, let the full model process tokens before s, then make the student predict the rest conditioned on the teacher’s cache. The loss combines cross-entropy with a KL term against the teacher’s top-64 token distribution plus a residual bucket.
for seq in corpus:
s = uniform(0, len(seq))
teacher_kv = full_model.prefill(seq[:s]) # frozen
student_logits = pruned_model.decode(
seq[s:], prefix_kv=teacher_kv)
loss = 0.1 * CE + 0.9 * KL_top64(student_logits, teacher_logits)
update(shared_expert, routers, norms) # ~0.42% of params
Only always-on components train: a newly added shared expert initialized to zero output (borrowing the LoRA trick), routers, and normalization layers. Routed experts and attention stay frozen. On Qwen3.6-35B-A3B that’s 0.42% of parameters; training uses ~100M tokens.
•
Training-free handoff recovers accuracy across criteria. At 50% pruning on Qwen3.6, SlimWise narrows the gap to the full model under all three criteria (routing mass, EAN, REAP), without touching the retained expert sets. The gain comes purely from giving the pruned decoder a full-model KV cache.
•
Benchmark scores hide length distortions. Under REAP, conventional pruning drops HumanEval+ median output from 2.3k to 0.3k tokens; on MATH-500 it balloons from 2.4k to 6.6k. The handoff restores HumanEval+ to 2.3k but leaves MATH-500 at ~5.8k. Distillation brings MATH-500 median back near 2.6k.
•
Distillation closes residual gaps. On Gemma 4-26B-A4B under EAN pruning, conventional pruning collapses tool-use benchmark Berkeley Function Calling Leaderboard to near zero because generations run to the 32,768-token budget. SlimWise alone doesn’t fix it, but SlimWise+Distill recovers accuracy to within a few points of baseline across math, coding, and tool use.
•
Throughput, measured in vLLM on 2×A100-80GB. Under PD-disaggregation at a 50 tokens/s per-user SLO, SlimWise reaches 1.40–1.81× full-model decode throughput at 50% pruning and 1.65–2.39× at 75% pruning. PD-colocated is slightly lower (1.36–1.72× and 1.57–2.28×), partly because router masking adds ~0.5 ms per decode step. Prefill throughput is unchanged by design.
•
Phase-swap ablation. Running pruned prefill with full decode (“Reverse”) helps math more; full prefill with pruned decode (SlimWise) helps coding more. Pruned decode is the main source of over-long math generations.
•
If you serve a large MoE under batched decoding and your traces show expert-weight traffic dominating decode step time, pruning only the decode phase is worth testing. The paper’s result is specific to Qwen3.6-35B-A3B and Gemma 4-26B-A4B with REAP/EAN/mass criteria; other backbones aren’t evaluated.
•
Measure generation length distributions, not just accuracy, when you prune or quantize a reasoning model. The paper shows pass-rate can hold while median output grows 2–3× or collapses to a few hundred tokens, both of which wreck throughput and latency SLOs in ways benchmark aggregates hide.
•
If you can afford a short training run (the authors report ~6 hours on 4 B200 GPUs updating under 1% of parameters), the distillation stage is where the length distortions actually get fixed. The training-free handoff alone is not enough for math-heavy workloads.
•
Under PD-colocated serving you keep the full expert pool resident, so you get the decode speedup but not the memory saving. If KV cache capacity is your binding constraint, disaggregation is where the larger gains come from.
•
SlimWise is implemented in vLLM; the paper doesn’t link a public artifact in the supplied text.
Prefill cost is unchanged, so end-to-end speedup depends on how decode-heavy your workload is. The distillation data mixture (open-perfectblend plus APIGen-MT traces) is specific; generalization to very different domains isn’t tested. Results are reported on two MoE backbones with k/m ratios in a particular range, and the extreme 75% pruning setting still shows MATH-500 90th-percentile output lengths well above baseline even after distillation. Finally, the paper doesn’t claim the handoff is universally superior: on EAN+Gemma 4, training-free SlimWise can underperform conventional pruning because the pruned decoder runs to the token budget, and only distillation rescues it.