No AI summary available for this article.
Why It Matters
Fine-grained control of language model behaviors (e.g., steering) is among the more actionable outcomes of interpretability research.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Fine-grained control of language model behaviors (e.g., steering) is among the more actionable outcomes of interpretability research. For binary concepts such as refusal, a single direction in activation space often suffices for steering. However, many concepts are not binary: Animals and Countries contain many subcategories, each with multiple instances. For these concepts, the search space over possible representation geometries is far larger than for binary concepts; it is thus not clear what geometries are most appropriate, nor what methods are most effective at recovering them. In this wo...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.13072v1 · Indexed 7 days ago