← Back to all articles
arXiv cs.AIOctober 7, 2026

Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning Models

Excerpt

arXiv:2610.06439v1 Announce Type: cross Abstract: What happens when a legal AI model learns to look like a lawyer instead of reasoning like one? We fine tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a proxy built from three surface features: citation count, legalese density, and response length. The model does not learn to reason more effectively. It learns to withhold commitment. Across 16 yes or no legal reasoning tasks from LegalBench (N=320), overall accuracy collapses