No AI summary available for this article.
Why It Matters
Production RL for language models lets the sampler fall behind the learner and repairs the resulting mismatch with a truncated importance weight.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Production RL for language models lets the sampler fall behind the learner and repairs the resulting mismatch with a truncated importance weight. We ask how long the sampler can go without a refresh under that correction, and find a cliff: on Qwen2.5-Math-1.5B and GSM8K, importance-corrected GRPO refreshed every 192 updates learns well for 180 steps and then degrades severely in all three data seeds before the refresh arrives. Published remedies for staleness act on the update; we act on the sampler instead. Decoupled cooling draws samples at temperature 0.8 while the learner, the reference mo...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.36953v1 · Indexed about 2 hours ago