← Back to all articles
arXiv cs.LGOctober 7, 2026

DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory

Excerpt

arXiv:2610.08553v1 Announce Type: new Abstract: Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outp