← Back to all articles
arXiv cs.LGOctober 2, 2026

Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It

Excerpt

arXiv:2610.00991v1 Announce Type: new Abstract: Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increas