No AI summary available for this article.
Why It Matters
Language models can answer from precomputed memory, a model's saved reading of a body of material, reused across requests instead of read again at each.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Language models can answer from precomputed memory, a model's saved reading of a body of material, reused across requests instead of read again at each. This paper maps where that practice preserves correctness and the conditions under which it fails. Across experiments on Llama-3.1-8B-Instruct using both saved key-value caches and trained compressions of them, precomputed memory degrades when assembled from separately prepared parts, stays current only through rebuilds costing a large fraction of full preparation in our measurements, and ignores corrections served beside it conditional on phr...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2608.30647v1 · Indexed about 2 hours ago