No AI summary available for this article.
Why It Matters
A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime monitoring as a means of completing ordinary tasks. We introduce EvasionBench, a benchmark of 50 diverse task-policy pairs in which completing the task requires an operation prohibited by a runtime monitor. Agents know that their tool calls are monitored and are prompted to continue working when they pause. Across our evaluations, best-of-3 evasion attempt rates reach up to 98% and succe...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.30217v1 · Indexed about 1 hour ago