STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
STEPQuant matches FP32 recurrent-state accuracy at about 6 bits while cutting serving memory.
STEPQuant is a spatial-temporal post-training quantization method for fixed-size delta-rule recurrent states used by linear attention. It assigns precision according to error magnitude and memory lifetime and jointly fits key-row and value-column scales. On Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct, it closely matches FP32-state accuracy at a nominal 6-bit budget and outperforms uniform INT8 at 4 bits. Integrated into SGLang, 6-bit STEPQuant provides over 5x recurrent-state compression and reduces total serving memory by up to 68.7%.
- Allocates precision by error magnitude and memory lifetime
- Jointly fits key-row and value-column quantization scales
- Matches FP32-state accuracy at a nominal 6-bit budget
- Beats uniform INT8 in its 4-bit configuration
- Cuts serving memory by up to 68.7% in SGLang
Full article191 words · extracted from arxiv.org · click to collapse
Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.38169