← Back to all articles
arXiv cs.LGOctober 2, 2026

Finetuning with Sampling: SFT Learns Better Than You Think

Excerpt

arXiv:2610.02140v1 Announce Type: new Abstract: Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL mu