arXiv cs.AIOctober 7, 2026
HuatuoGPT-3: RL-Only Domain Adaptation from Base Models
Excerpt
arXiv:2610.05966v1 Announce Type: cross Abstract: Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient cold start, it may reduce exploration diversity and introduces additional complexity through multi-stage optimization. These limitations motivate RL-only adaptation. However, pure on-policy RL suffers from a cold-start problem, while mixed-policy RL still falls short: informative to