← Back to all articles
arXiv cs.LGOctober 2, 2026

Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation

Excerpt

arXiv:2607.14895v2 Announce Type: replace Abstract: Reasoning language models (RLMs) demonstrate impressive performance by leveraging test-time compute in the form of reasoning tokens. However, this behavior makes adapting RLMs to new domains challenging and expensive. The reason is that further training can disturb the learned behavior and degrade model performance. This makes it difficult to leverage supervised fine-tuning data with human-written solutions: although it contains high-quality an