arXiv cs.LGOctober 2, 2026
Finetuning with Sampling: SFT Learns Better Than You Think
Excerpt
arXiv:2610.02140v1 Announce Type: new Abstract: Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL mu