arXiv cs.LGOctober 1, 2026
SOLAR: SVD-Optimized Lifelong Attention for Recommendation
Excerpt
arXiv:2603.02561v2 Announce Type: replace-cross Abstract: Attention mechanism remains the defining operator in Transformers since it provides expressive global credit assignment, yet its quadratic cost in sequence length N makes long-context modeling expensive and often forces truncation or other heuristics. Linear attention reduces complexity to O(Nd^2) by reordering computation through kernel feature maps, but this reformulation drops the softmax mechanism and shifts the attention score distri