No AI summary available for this article.
Why It Matters
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit sp...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2610.02185v1 · Indexed about 1 hour ago