TY - GEN
T1 - Soft-robust actor-critic policy-gradient
AU - Derman, Esther
AU - Mankowitz, Daniel J.
AU - Mann, Timothy A.
AU - Mannor, Shie
N1 - Funding Information: This work was partially funded by the Israel Science Foundation under contract 1380/16 and by the European Community’s Seventh Framework Programme (FP7/2007-2013) under grant agreement 306638 (SUPREL). Publisher Copyright: © 2018 by Association For Uncertainty in Artificial Intelligence (AUAI) All rights reserved.
PY - 2018
Y1 - 2018
N2 - Robust Reinforcement Learning aims to derive an optimal behavior that accounts for model uncertainty in dynamical systems. However, previous studies have shown that by considering the worst case scenario, robust policies can be overly conservative. Our soft-robust framework is an attempt to overcome this issue. In this paper, we present a novel Soft-Robust Actor-Critic algorithm (SR-AC). It learns an optimal policy with respect to a distribution over an uncertainty set and stays robust to model uncertainty but avoids the conservativeness of robust strategies. We show the convergence of SR-AC and test the efficiency of our approach on different domains by comparing it against regular learning methods and their robust formulations.
AB - Robust Reinforcement Learning aims to derive an optimal behavior that accounts for model uncertainty in dynamical systems. However, previous studies have shown that by considering the worst case scenario, robust policies can be overly conservative. Our soft-robust framework is an attempt to overcome this issue. In this paper, we present a novel Soft-Robust Actor-Critic algorithm (SR-AC). It learns an optimal policy with respect to a distribution over an uncertainty set and stays robust to model uncertainty but avoids the conservativeness of robust strategies. We show the convergence of SR-AC and test the efficiency of our approach on different domains by comparing it against regular learning methods and their robust formulations.
UR - https://www.scopus.com/pages/publications/85059364037
M3 - Conference contribution
T3 - 34th Conference on Uncertainty in Artificial Intelligence 2018, UAI 2018
SP - 208
EP - 218
BT - Uncertainty in Artificial Intelligence - Proceedings of the 34th Conference, UAI 2018
A2 - Globerson, Amir
A2 - Silva, Ricardo
PB - Association For Uncertainty in Artificial Intelligence (AUAI)
T2 - 34th Conference on Uncertainty in Artificial Intelligence, UAI 2018
Y2 - 6 August 2018 through 10 August 2018
ER -