LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
LeapQuant nearly matches FP32 quality with training-free 8-bit recurrent-state quantization.
LeapQuant is a training-free method for quantizing recurrent states in hybrid LLMs that use linear attention, including Gated DeltaNet and Kimi Delta Attention. Per-window quantization reduces accumulated rounding error, while a few high-precision compensator tokens retain the largest outliers before the residual is smoothed and quantized. Across the Qwen, Kimi, and GLM families, it reports accuracy comparable to an FP32 baseline. It achieves average kernel speedups of 2.05-3.70x and 1.47x end-to-end inference speedup on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.
44