No AI summary available for this article.
Why It Matters
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly auditable. FTA contains 100 tasks with deterministic failure traces spanning five failure families, a neutral control, and four user-pressure conditions, and...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.35732v1 · Indexed 39 minutes ago