No AI summary available for this article.
Why It Matters
Multi-turn tool-use failures can hinge on a single model call, yet reward variation alone does not reveal which call would benefit from training.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Multi-turn tool-use failures can hinge on a single model call, yet reward variation alone does not reveal which call would benefit from training. When rewards depend on later interactions, their variation can reflect downstream randomness rather than differences between the current actions. We introduce Critical-State RL to identify trainable states in multi-turn interactions. Given task-defined candidate calls and local rewards, the method assesses whether each reward captures the action's effect on task success and whether improvement over a reference policy is possible. It then uses nested...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.24985v1 · Indexed about 1 hour ago