No AI summary available for this article.
Why It Matters
A small trainable advisor can steer a frozen language-model executor using natural-language advice.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the mo...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.38142v1 · Indexed about 1 hour ago