No AI summary available for this article.
Why It Matters
We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set. In the offline setting, where the reward function is known, we show that convexity and cyclic-shift invariance make a stationary solution optimal when $w\mid T$ and within an additive $O(w)$ gap otherwise. In the online setting, we show that geometric structure alone is insufficient for learning, and sublinear regret can be impossible. We introduce a transition diameter $τ$ that quantifies feasible reachability and develop a rare-switching OFUL algo...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2610.08745v1 · Indexed about 1 hour ago