← Back to all articles
arXiv cs.AIAugust 18, 2026

Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning

Excerpt

arXiv:2608.14963v1 Announce Type: cross Abstract: Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command. However, the local mapping from command and state to action remains opaque. We propose command-space counterfactual explanations for PCNs: given a fixed state, original command, and foil action, we search, in a black-box setting, for a minimally changed desired-return command under which the sam