← Back to all articles
arXiv cs.LGOctober 1, 2026

Evaluating Agents Across Runtime Contracts: When Mismatch Costs Efficiency or Quality

Excerpt

arXiv:2603.01209v3 Announce Type: replace-cross Abstract: In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python variables between actions; stateless runtimes clear them without resetting task progress. Training traces demonstrate task-solving strategies and runtime-specific ways to store and recover intermediate results. We study runtime transfer: whether agents trained under one contract remain effective