No AI summary available for this article.
Why It Matters
Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution. The resulting Fisher quadratic measures distributional sensitivity and admits an analytic gradient computed through matrix--vector products with the frozen language-model head. We derive an exact decompositi...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.30218v1 · Indexed about 1 hour ago