No AI summary available for this article.
Why It Matters
Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We show that persistent KV-cache reuse with selective recomputation can be history-dependent: in our rolling-agent workload, an unchanged prompt can produce different answers depending on the requests processed before it. At a matched 5% recomputation budget, document-aligned recomputation reduces answer variation across request orders from 6...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2610.05833v1 · Indexed 42 minutes ago