No AI summary available for this article.
Why It Matters
Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study this evolution as a functional geometry, decomposing learned updates into parallel and perpendicular components. Across pretrained models, we find substantial parallel components beyond the residual identity path. We then apply the decomposition in two spaces: to attention and MLP updates relative to the hidden state, and to attention value aggregation relative to the current token's value. Targeted edits reveal a strongly space-dependent asymmetry:...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.15975v1 · Indexed about 1 hour ago