No AI summary available for this article.
Why It Matters
Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cross-chunk attention information, reducing answer quality. Selective recomputation methods recover the missing cross-chunk context by rerunning the target LLM on selected tokens, incurring substantial online computation. We introduce CacheRepair, a lightweight network that learns...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.35139v1 · Indexed about 1 hour ago