No AI summary available for this article.
Why It Matters
Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2610.08743v1 · Indexed about 1 hour ago