arXiv cs.LGOctober 2, 2026
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Excerpt
arXiv:2610.02199v1 Announce Type: new Abstract: Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead t