arXiv cs.LGOctober 1, 2026
PhantomEnvironments: Training LLM Agents in Fictional Worlds
Excerpt
arXiv:2609.40221v1 Announce Type: new Abstract: Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose g