No AI summary available for this article.
Why It Matters
Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a model minimally edits an input so that its own prediction changes. Across sentiment analysis and natural language inference, we evaluate ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales. Our results show that model scale is the strongest determinant of explanation qual...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.17119v1 · Indexed about 1 hour ago