← Back to all articles
arXiv cs.LGOctober 1, 2026

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Excerpt

arXiv:2609.39102v1 Announce Type: cross Abstract: Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive