No AI summary available for this article.
Why It Matters
Synthetic data is increasingly used to train large language models (LLMs), yet its security implications remain poorly understood.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Synthetic data is increasingly used to train large language models (LLMs), yet its security implications remain poorly understood. Prior work on subliminal learning suggests that models can inherit behavioral traits from seemingly unrelated training data. In this work, we investigate whether such mechanisms can be exploited to inject targeted social biases into aligned models through semantically benign synthetic data. We construct a pipeline in which a misaligned teacher model generates filtered synthetic datasets across domains such as creative writing and code generation, which are then use...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2608.30619v1 · Indexed about 2 hours ago