No AI summary available for this article.
Why It Matters
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain pos...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.20779v1 · Indexed about 1 hour ago