No AI summary available for this article.
Why It Matters
We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE -- referred to as RE(S) -- that updates the rollout distribution once every $S \ge 1$ gradient steps.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE -- referred to as RE(S) -- that updates the rollout distribution once every $S \ge 1$ gradient steps. Prior work in bandits and reinforcement learning has developed rich theory for policy gradient methods, and on-policy sampling (i.e., a small $S$, ideally $1$) is often viewed as crucial to their success; yet in prominent application like post-training large language models, reward-guided self-training has proved to be effective ev...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.36945v1 · Indexed about 2 hours ago