← Back to all articles
arXiv cs.AIOctober 7, 2026

Selecting Long-Horizon Trajectories for Reliable and Efficient Terminal-Agent Training

Excerpt

arXiv:2610.05831v1 Announce Type: cross Abstract: Terminal agents are commonly trained by imitating long teacher trajectories, yet how much of each trajectory to supervise remains unexplored. We study the \emph{supervision horizon}, the number of trajectory tokens retained for training, and show that it is a key design axis for reliability and cost. Reliability improves with longer horizons but saturates: on Terminal-Bench, a 12K-token horizon solves more tasks than 16K ($29\pm0.7$ vs.\ $26\pm0.