No AI summary available for this article.
Why It Matters
Regularization-based methods have become a standard approach for training Deep Reinforcement Learning policies against adversarial input perturbations.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Regularization-based methods have become a standard approach for training Deep Reinforcement Learning policies against adversarial input perturbations. In this paper, we unify these methods by deriving new upper bounds on the performance gap between the nominal and worst-case policies. Each upper bound is expressed as an existing regularization objective plus a KL-divergence penalty between the nominal and worst-case policies, which further explains why adding a KL penalty improves robustness in practice. Building on these bounds, we formulate robust training as a constrained optimization prob...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.13050v1 · Indexed 7 days ago