No AI summary available for this article.
Why It Matters
Automated reference-based evaluation methods play a critical role in assessing natural language generation systems.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Automated reference-based evaluation methods play a critical role in assessing natural language generation systems. Existing meta-evaluation primarily measures agreement with human judgments or benchmark labels, providing limited insight into evaluator behavior under controlled conditions. We introduce behavioral correctness assumptions, a complementary framework for evaluating reference-based automatic evaluation methods. We define a taxonomy of correctness-preserving and correctness-altering assumptions and operationalize them through controlled response transformations that specify expected...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.05289v1 · Indexed 5 days ago