arXiv cs.AIOctober 2, 2026
PhGPO: Pheromone-Guided Policy Optimization for Long-Horizon Tool Planning
Excerpt
arXiv:2602.13691v2 Announce Type: replace Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated strong capabilities in executing complex tasks through tool use. However, long-horizon multi-step tool planning is challenging, because the exploration space suffers from a combinatorial explosion. In this scenario, even when a correct tool-use path is found, it is usually considered an immediate reward for current training, which would not provide any reusable informati