arXiv cs.LGOctober 1, 2026
OPSRD: On-Policy Self-Role Distillation
Excerpt
arXiv:2609.39884v1 Announce Type: cross Abstract: Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, whi