No AI summary available for this article.
Why It Matters
Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperfor...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.35748v1 · Indexed 39 minutes ago