arXiv cs.CLSeptember 21, 2026
Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms
Excerpt
arXiv:2608.27409v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models, but covering multiple capabilities often involves training separate domain experts and subsequently consolidating them. We organize three fusion paradigms by the artifacts they reuse: Merge combines expert task vectors, Mix RL pools their datasets, and multi-teacher on-policy distillation (MOPD) uses both. Because they have largely bee