No AI summary available for this article.
Why It Matters
Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, re...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2608.20281v1 · Indexed 8 days ago