arXiv cs.AIOctober 7, 2026
The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning
Excerpt
arXiv:2610.06446v1 Announce Type: cross Abstract: Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on