← Back to all articles
arXiv cs.AIOctober 7, 2026

Clean: Second-order LLM Training at Linear Memory Cost via Nystr\"om Sketching

Excerpt

arXiv:2610.04204v1 Announce Type: cross Abstract: Training large language models (LLMs) entails a fundamental trade-off: memory-efficient optimizers such as Adam discard cross-parameter curvature, whereas full-curvature methods such as SOAP can accelerate convergence at prohibitive memory costs. We introduce Clean, a memory-efficient and full-curvature optimizer designed to resolve this bottleneck. Clean leverages the randomized Nystrom method to accurately approximate the left and right precond