No AI summary available for this article.
Why It Matters
Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the spectral properties of linear attention state dynamics. In this work, we study long-context extension of Gated DeltaNet (GDN) from a spectral perspective of transition matrix and identify two essential factors governing long-range information retrieval: (1) a sufficiently broad slow spectral band aligned with the target depend...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.14320v1 · Indexed about 1 hour ago