WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
WUSH-KV uses data-adaptive transforms for low-bit KV-cache quantization and matches or beats OSCAR at 2 bits.
WUSH-KV applies the WUSH data-adaptive transform, built from second-order statistics of both factors in a matrix product, to low-bit KV-cache quantization for long-context inference. Calibration data produces separate key and value transforms: the value transform is folded into model weights and the key transform is applied after RoPE, and both can pair with clipped quantizers. For the QuEST INT quantizer, the authors show the WUSH transform is near-optimal under mild assumptions, cuts layerwise reconstruction error, and records the lowest end-to-end perplexity among tested transforms. Integrated into SGLang with OSCAR-style percentile-clipped affine quantization, 2-bit WUSH-KV matches or beats the OSCAR transform across evaluated models and downstream tasks.
- WUSH-KV builds separate data-adaptive transforms for keys and values.
- Value transform is folded into weights; key transform is applied after RoPE.
- With QuEST INT, WUSH is shown near-optimal and yields the lowest perplexity tested.
- At 2-bit in SGLang, it matches or beats the OSCAR transform on evaluated tasks.
Full article153 words · extracted from arxiv.org · click to collapse
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.38121