No AI summary available for this article.
Why It Matters
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens are processed. Because cached states are computed in the context of earlier tokens, they may carry information from beyond the current window and transmit it to later states. This study presents a series of experiments using five open-weight models spanning Qwen, Llama, Mistral, and Muse Glimmer. We...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.34049v1 · Indexed 42 minutes ago