← Back to all articles
arXiv cs.AIOctober 7, 2026

CORE-RL: Confidence-Oriented Reliability Evaluation of Black-Box Reinforcement Learning Policies

Excerpt

arXiv:2610.04418v1 Announce Type: new Abstract: The deployment of Reinforcement Learning (RL) agents in critical domains must be preceded with a pipeline to evaluate the alignment of the RL agent with complex multi-objective specifications and robustness under real-world environmental drift. However, to protect intellectual property, the RL agent may be delivered for evaluation as opaque executable or remote API, which makes traditional evaluation techniques based on the internals of the policie