arXiv cs.LGOctober 7, 2026
DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory
Excerpt
arXiv:2610.08553v1 Announce Type: new Abstract: Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outp