arXiv cs.LGOctober 7, 2026
FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
Excerpt
arXiv:2610.07898v1 Announce Type: new Abstract: Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group. However, this train