SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
SparseDecoding prunes LLMs using decoding activations, speeding generation up to 1.48x.
SparseDecoding is a decoding-aware, training-free pruning method for memory-bound LLM inference. It builds calibration matrices from layer-wise activations collected during dense-model autoregressive generation, excluding prefill, so the pruning objective matches decoding rather than pre-collected natural text. It also adds an N:M sparse matrix-vector kernel with bitmask indexing and fixed-step traversal. On Llama-3.1-8B, Llama-3.3-70B, and Qwen3-14B/32B it beats fixed-text calibration on long-form generation and reaches up to 1.48x end-to-end decoding speedup on A100 GPUs.
- Hessian pruning on natural text mismatches decoding-time activations.
- Calibration uses layer activations from autoregressive generation, excluding prefill.
- An N:M sparse matrix-vector kernel uses bitmask indexing.
- Up to 1.48x decoding speedup on A100 for Llama and Qwen3 models.
Full article248 words · extracted from huggingface.co · click to collapse
The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding. Nevertheless, typical methods in this line compute the Hessian using pre-collected natural sequences, whereas the model is fed self-generated tokens during decoding, creating a distribution shift between the two sequences. The Hessian calculated on the natural sequence is different from that calculated on the generated sequence. We observe that this discrepancy causes the activation distribution during generation to deviate from that used for pruning, further hurting the pruned model performance. Moreover, most existing LLM pruning methods that bring actual speedup primarily target the sparse matrix-matrix (SpMM) multiplication, providing limited support for the sparse matrix-vector (SpMV) operations, which dominate decoding. To solve these problems, we introduce SparseDecoding, a principled decoding-aware pruning framework tailored for accurate and efficient LLM decoding. Specifically, at the algorithmic axis, SparseDecoding constructs calibration matrices from layer-wise activations collected during the dense-model autoregressive generation, excluding prefill, thereby aligning the pruning objective with the decoding activations. At the system axis, we develop an optimized N:M sparse matrix-vector kernel with bitmask indexing and fixed-step traversal. Substantial empirical results on representative LLMs (Llama-3.1-8B, Llama-3.3-70B, Qwen3-14B / 32B) demonstrate that our method consistently outperforms standard fixed-text calibration on the long-form generation benchmarks while achieving up to 1.48x end-to-end wall-clock decoding speedup on A100 GPUs.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.12327