← Back to all articles
arXiv cs.CLAugust 19, 2026

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

Excerpt

arXiv:2608.17804v1 Announce Type: cross Abstract: Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing fo