No AI summary available for this article.
Why It Matters
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure \emph{state coverage}, the fraction...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.04172v1 · Indexed about 2 hours ago