Skip to main navigation Skip to search Skip to main content

Mean-variance optimization in Markov decision processes

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove that the complexity of computing a policy that maximizes the mean reward under a variance constraint is NP-hard for some cases, and strongly NP-hard for others. We finally offer pseudopolynomial exact and approximation algorithms.

Original languageEnglish GB
Title of host publicationProceedings of the 28th International Conference on Machine Learning, ICML 2011
Pages177-184
Number of pages8
StatePublished - 2011
Event28th International Conference on Machine Learning, ICML 2011 - Bellevue, WA, United States
Duration: 28 Jun 20112 Jul 2011

Publication series

NameProceedings of the 28th International Conference on Machine Learning, ICML 2011

Conference

Conference28th International Conference on Machine Learning, ICML 2011
Country/TerritoryUnited States
CityBellevue, WA
Period28/06/112/07/11

ASJC Scopus subject areas

  • Computer Science Applications
  • Human-Computer Interaction
  • Education

Fingerprint

Dive into the research topics of 'Mean-variance optimization in Markov decision processes'. Together they form a unique fingerprint.

Cite this