WUSH-KV Applies Data-Adaptive Transforms to KV Cache Quantization
WUSH-KV uses calibration-based key and value transforms so 2-bit KV-cache quantization matches or beats OSCAR.
WUSH-KV applies the WUSH data-adaptive transform, derived from second-order statistics of both factors in a matrix product, to low-bit key-value cache quantization for long-context inference. Calibration data yields separate key and value transforms: the value transform is folded into model weights and the key transform is applied after RoPE, and both can pair with clipped quantizers. For the QuEST INT quantizer, the authors say the transform is near-optimal under mild assumptions, reduces layerwise reconstruction error, and records the lowest end-to-end perplexity among transforms they tested. Integrated into SGLang with OSCAR-style percentile-clipped affine quantization, 2-bit WUSH-KV matches or beats the OSCAR transform across evaluated models and downstream tasks. The reports do not disagree; the Sep 29 arXiv account is more specific about long-context use, mild assumptions, and reconstruction error than the Sep 28 Hugging Face note.
- WUSH-KV applies a data-adaptive transform built from second-order statistics of both factors in a matrix product to low-bit KV-cache quantization.
- Calibration data produces separate key and value transforms; the value transform is folded into model weights and the key transform is applied after RoPE.
- Both transforms can be paired with clipped quantizers.
- For QuEST INT, the authors say WUSH is near-optimal under mild assumptions, cuts layerwise reconstruction error, and has the lowest end-to-end perplexity among tested transforms.
- Integrated into SGLang with OSCAR-style percentile-clipped affine quantization, 2-bit WUSH-KV matches or beats the OSCAR transform across evaluated models and downstream tasks.
- The Sep 28 and Sep 29 reports agree; only the later account states long-context use, mild optimality assumptions, and lower reconstruction error.
Coverage timelineoldest first · each row is one article
- · 3d agoWUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
Hugging Face daily papers· 50
WUSH-KV quantizes KV caches with data-adaptive transforms and matches or beats OSCAR at 2 bits.
- · 2d agoWUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
arXiv cs.AI / cs.LG / cs.CL· 48
WUSH-KV uses data-adaptive transforms for low-bit KV-cache quantization and matches or beats OSCAR at 2 bits.