arXiv cs.AIOctober 7, 2026
Direction-Conditioned Policies for Online Goal-Conditioned Reinforcement Learning
Excerpt
arXiv:2610.05087v1 Announce Type: cross Abstract: Contrastive Reinforcement Learning (CRL) learns representations that estimate goal reachability, yet its policy remains conditioned on raw goals and therefore does not directly exploit the geometry encoded by its critic. We introduce Direction-Conditioned Policies (DCP), a method built around a small modification to CRL: DCP selects previously visited states as waypoints during online training and conditions the policy on their direction and dist