← Back to all articles
arXiv cs.AIOctober 7, 2026

The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning

Excerpt

arXiv:2610.06446v1 Announce Type: cross Abstract: Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on