arXiv cs.AIOctober 7, 2026
Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers
Excerpt
arXiv:2610.06204v1 Announce Type: new Abstract: LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective parameters) with GRPO on a programmatic utility reward for bilateral multi-issue bargaining, and evaluate every arm on the same 1,152