← Back to all articles
arXiv cs.LGOctober 1, 2026

ASH: Agents that Self-Hone in Long-Horizon Worlds

Excerpt

arXiv:2605.14211v4 Announce Type: replace-cross Abstract: Long-horizon visuomotor tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrations, neither of which scales. We introduce ASH, an agentic system that learns a long-horizon policy from unlabeled, noisy internet video, without reward shaping or expert annotation. ASH follows a self-improvement loop; when it gets stuck, ASH learns an Inverse Dynamics Model (IDM) from its ow