← Back to all articles
OpenAI BlogMarch 20, 2018

Variance reduction for policy gradient with action-dependent factorized baselines