No AI summary available for this article.
Why It Matters
When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score. That score cannot tell whether improvement was attainable, which component lost value, or what the agent's own checks certify. We introduce a claim-specific verification audit for modular agents that plan, act, check and refine. Instead of scoring the agent, the audit scores the evidence: each conclusion is recorded with the evidence behind it, one of four verdicts (supported, unsupported, unresolved or not evaluated) and the bound...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2610.01348v1 · Indexed about 2 hours ago