← Back to all articles
arXiv cs.LGOctober 1, 2026

Low-Discrepancy Dither for Quantized Recurrent State Caches

Excerpt

arXiv:2609.39185v1 Announce Type: new Abstract: Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed back into the next update and can accumulate over long generations. Production systems round the state stochastically; we ask which rounding rule such caches should use. We find that a deterministic golden-ratio Weyl dither