No AI summary available for this article.
Why It Matters
On-policy distillation trains a language model on its own generations while a teacher scores them token by token.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were pro...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2608.25936v1 · Indexed 4 days ago