No AI summary available for this article.
Why It Matters
Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed back into the next update and can accumulate over long generations. Production systems round the state stochastically; we ask which rounding rule such caches should use. We find that a deterministic golden-ratio Weyl dither, which needs no random numbers, consistently brings the quantized model closer to the full-precision one than stochastic rounding, across pure and...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.39185v1 · Indexed about 1 hour ago