arXiv cs.AIOctober 7, 2026
HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models
Excerpt
arXiv:2610.05739v1 Announce Type: cross Abstract: Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we identify severe long-range forgetting in Gated DeltaNet (GDN), where information from distant but relevant scenes is progressivel