No AI summary available for this article.
Why It Matters
Manipulating next-token probabilities during generation can bypass the safety alignment of large language models.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.36956v1 · Indexed about 2 hours ago