← Back to all articles
arXiv cs.AIOctober 7, 2026

Adaptive Expert Guidance for Efficient On-Policy Reinforcement Learning

Excerpt

arXiv:2610.06019v1 Announce Type: cross Abstract: With massively parallel simulation, on-policy Reinforcement Learning methods such as PPO have become standard in many domains. However, learning from scratch is sample-inefficient and fails to exploit the potential existence of a suboptimal expert, such as a heuristic, a model-based controller, or a policy trained on a related task. Such an expert is often available and can guide early training, but its sub-optimality limits final performance. Th