LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
LeapQuant nearly matches FP32 quality with training-free 8-bit recurrent-state quantization.
LeapQuant is a training-free method for quantizing recurrent states in hybrid LLMs that use linear attention, including Gated DeltaNet and Kimi Delta Attention. Per-window quantization reduces accumulated rounding error, while a few high-precision compensator tokens retain the largest outliers before the residual is smoothed and quantized. Across the Qwen, Kimi, and GLM families, it reports accuracy comparable to an FP32 baseline. It achieves average kernel speedups of 2.05-3.70x and 1.47x end-to-end inference speedup on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.
- Training-free 8-bit quantization for recurrent linear-attention states
- Per-window quantization limits accumulated rounding error
- High-precision compensator tokens preserve large state outliers
- Reports 2.05-3.70x kernel and 1.47x end-to-end speedups
Full article239 words · extracted from arxiv.org · click to collapse
Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70$\times$ at the kernel level and 1.47$\times$ for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.38166