SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving
SlimWise prunes MoE experts only at decode, lifting Qwen3.6 throughput up to 1.81x in vLLM.
SlimWise is a training-light Mixture-of-Experts serving framework that runs prefill on the full model and decode on a pruned expert pool, reusing the prefill KV cache without conversion. Across two MoE backbones and three pruning criteria, the handoff narrows accuracy gaps versus the full model, though benchmark scores can hide changes in generation length. A low-cost distillation stage trains the decoder to continue from full-model KV caches while updating few parameters. Implemented in vLLM, SlimWise improves decode throughput by up to 1.81x on Qwen3.6-35B-A3B at 50% expert pruning with minimal accuracy loss.
- Prefill uses the full MoE; decode uses a pruned expert pool.
- Pruned decode reuses the prefill KV cache without conversion.
- A small distillation stage updates only a subset of decoder parameters.
- vLLM implementation supports disaggregated and colocated prefill-decode.
- On Qwen3.6-35B-A3B, 50% pruning yields up to 1.81x decode throughput.
Full article181 words · extracted from huggingface.co · click to collapse
Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.34117