No AI summary available for this article.
Why It Matters
Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families....
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.04166v1 · Indexed about 2 hours ago