No AI summary available for this article.
Why It Matters
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy disti...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.34036v1 · Indexed 41 minutes ago