← Back to all articles
arXiv cs.LGOctober 2, 2026

Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation

Excerpt

arXiv:2610.00729v1 Announce Type: new Abstract: This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neural network policies often struggle to adapt to a new environment and require a considerable amount of samples for successful transfer. Instead, we propose a novel reward-based policy only conditioned on rewards and action