Reddit r/MachineLearningOctober 6, 2026
Transformers vs RNNs vs SSMs: Where Does Memory Actually Live? [D]
Excerpt
Someone who has always loved looking at the space between different AI techniques, this time I went a little deeper into the memory trade-offs between RNNs, Transformers and SSMs. I found it interesting because once you start looking at these architectures through the lens of working memory, a lot of the differences become easier to understand. Where does the memory actually live? Is it a compact recurrent state, a growing KV cache, or something closer to the network itself? RNNs keep memory in