arXiv cs.AIOctober 7, 2026
Recursive Video In-Context Learning for Agentic Robot
Excerpt
arXiv:2610.06843v1 Announce Type: cross Abstract: LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames