← Back to all articles
arXiv cs.LGAugust 18, 2026

MoRFI: Monotonic Sparse Autoencoder Feature Identification

Excerpt

arXiv:2604.26866v2 Announce Type: replace-cross Abstract: Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of post-training often introduce new facts outwith the parametric knowledge, giving rise to hallucinations. While it has been demonstrated that supervised fine-tuning (SFT) on new knowledge may exacerbate the problem, the underlying mechanisms are still poorly understood. We conduct a control