← Back to all articles
arXiv cs.LGOctober 2, 2026

The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning

Excerpt

arXiv:2610.00332v1 Announce Type: new Abstract: Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at correct final answers through flawed intermediate logic, while regularizing with soft divergence penalties against a teacher (e.g., KL-based distillati