← Back to all articles
arXiv cs.LGOctober 7, 2026

What pass@k Cannot Measure: Evaluating Diversity and Capability Retention after Post-Training

Excerpt

arXiv:2610.07405v1 Announce Type: new Abstract: pass@$k$, the fraction of problems a model solves within $k$ sampled attempts, is the field's default protocol for deciding whether reinforcement-learning (RL) post-training on verifiable rewards improved a model. At the population level, pass@$k$ depends only on a problem's probability of a correct sample, with no term for how it is distributed across outputs. We show this gap is not academic. Training Qwen2.5-1.5B-Instruct on grade-school math wi