← Back to all articles
arXiv cs.AIOctober 7, 2026

CIPO: Counterfactual Imagination Policy Optimization for Adaptive Tool Granularity Selection

Excerpt

arXiv:2610.04991v1 Announce Type: new Abstract: Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite "Skills" can shorten tool-use trajectories and reduce repeated low-level decisions. However, when atomic tools and composite skills coexist, skill use becomes a policy problem: the agent must decide whether the current state requires atomic fi