Skip to main navigation Skip to search Skip to main content

Soft-robust actor-critic policy-gradient

  • Esther Derman
  • , Daniel J. Mankowitz
  • , Timothy A. Mann
  • , Shie Mannor

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Robust Reinforcement Learning aims to derive an optimal behavior that accounts for model uncertainty in dynamical systems. However, previous studies have shown that by considering the worst case scenario, robust policies can be overly conservative. Our soft-robust framework is an attempt to overcome this issue. In this paper, we present a novel Soft-Robust Actor-Critic algorithm (SR-AC). It learns an optimal policy with respect to a distribution over an uncertainty set and stays robust to model uncertainty but avoids the conservativeness of robust strategies. We show the convergence of SR-AC and test the efficiency of our approach on different domains by comparing it against regular learning methods and their robust formulations.

Original languageEnglish
Title of host publicationUncertainty in Artificial Intelligence - Proceedings of the 34th Conference, UAI 2018
EditorsAmir Globerson, Ricardo Silva
PublisherAssociation For Uncertainty in Artificial Intelligence (AUAI)
Pages208-218
Number of pages11
ISBN (Electronic)9781510871601
StatePublished - 2018
Event34th Conference on Uncertainty in Artificial Intelligence, UAI 2018 - Monterey, United States
Duration: 6 Aug 201810 Aug 2018

Publication series

Name34th Conference on Uncertainty in Artificial Intelligence 2018, UAI 2018
Volume1

Conference

Conference34th Conference on Uncertainty in Artificial Intelligence, UAI 2018
Country/TerritoryUnited States
CityMonterey
Period6/08/1810/08/18

ASJC Scopus subject areas

  • Artificial Intelligence

Fingerprint

Dive into the research topics of 'Soft-robust actor-critic policy-gradient'. Together they form a unique fingerprint.

Cite this