No AI summary available for this article.
Why It Matters
Language models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Language models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts. We investigate whether token-indexed memory can replace the dedicated value projection when complemented by contextual information. We propose Memory Attention (MA), which forms values by combining layer-specific token memory with contextual keys. The memory supplies token-specific representations, while the keys preserve context dependence. At inference, normalization can be folded into the memory tables, reducing value construction to lookup and...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.28399v1 · Indexed about 2 hours ago