No AI summary available for this article.
Why It Matters
We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs).
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs). We propose a simple analysis method: For each neuron, we compute the cosine similarities between its input (reading) and output (writing) weight vectors. In this scheme, a strong negative cosine similarity indicates the neuron weakens the direction it detects in the residual stream, so we call this a weakening neuron. This allows us to gain a number of novel insights. First, we show that nine different LLMs have similar patterns: weakening neurons appear mostly in late layers whereas their counte...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.18612v1 · Indexed about 1 hour ago