← Back to all articles
arXiv cs.LGOctober 7, 2026

Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers

Excerpt

arXiv:2608.07436v2 Announce Type: replace-cross Abstract: Muon-trained modular-arithmetic transformers can lose accuracy while retaining linearly decodable task information. Adjacent swaps localize five captured unnormalized failures to AdamW readout updates. Multiplying the actual readout displacement by the large feature mean produces a class-dependent logit offset shared across inputs that nearly reproduces each failure. Training-only decoders recover 98.20-100% held-out accuracy. Correcting