No AI summary available for this article.
Why It Matters
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.38121v1 · Indexed about 1 hour ago