Two methods quantize linear-attention states near FP32 quality
LeapQuant and STEPQuant lower recurrent-state precision for linear-attention LLMs while staying near FP32 accuracy and cutting compute or memory.
Two arXiv papers dated 2026-09-29 propose ways to quantize the fixed recurrent states used by linear-attention language models while staying close to full-precision quality. LeapQuant is a training-free 8-bit method for hybrid models including Gated DeltaNet and Kimi Delta Attention: per-window quantization limits accumulated rounding error, and a few high-precision compensator tokens retain the largest outliers before the residual is smoothed and quantized. Across the Qwen, Kimi, and GLM families it reports accuracy comparable to an FP32 baseline, with average kernel speedups of 2.05-3.70x and 1.47x end-to-end inference speedup on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs. STEPQuant is a spatial-temporal post-training method that assigns precision according to error magnitude and memory lifetime and jointly fits key-row and value-column scales. On Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct it closely matches FP32-state accuracy at a nominal 6-bit budget, outperforms uniform INT8 at 4 bits, and in SGLang yields over 5x recurrent-state compression and up to 68.7% less total serving memory. The sources cover distinct techniques and do not report a head-to-head comparison.
- Both papers were posted to arXiv (cs.AI, cs.LG, cs.CL) on 2026-09-29 and target quantization of fixed recurrent states in linear-attention LLMs.
- LeapQuant is training-free 8-bit quantization for hybrid models including Gated DeltaNet and Kimi Delta Attention, using per-window quantization and a few high-precision compensator tokens for large outliers.
- Across Qwen, Kimi, and GLM, LeapQuant reports accuracy comparable to FP32, average kernel speedups of 2.05-3.70x, and 1.47x end-to-end inference speedup on NVIDIA B200, RTX PRO 6000, and RTX 5090.
- STEPQuant is spatial-temporal post-training quantization that allocates precision by error magnitude and memory lifetime and jointly fits key-row and value-column scales.
- On Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct, STEPQuant closely matches FP32-state accuracy at a nominal 6-bit budget and outperforms uniform INT8 at 4 bits.
- In SGLang, 6-bit STEPQuant provides over 5x recurrent-state compression and reduces total serving memory by up to 68.7%.
- The reports describe different methods and metrics and do not directly compare LeapQuant with STEPQuant.
Coverage timelineoldest first · each row is one article
- · 2d agoLeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
arXiv cs.AI / cs.LG / cs.CL· 44
LeapQuant nearly matches FP32 quality with training-free 8-bit recurrent-state quantization.
- · 2d agoSTEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
arXiv cs.AI / cs.LG / cs.CL· 44
STEPQuant matches FP32 recurrent-state accuracy at about 6 bits while cutting serving memory.