TY - JOUR
T1 - Distributionally robust markov decision processes
AU - Xu, Huan
AU - Mannor, Shie
N1 - Funding Information: The authors gratefully acknowledge the assistance, support, and direct contributions of Mr. Thomas A. Furness, III, Chief, Visual Display Systems Branch, Human Engineering Division, Aerospace Medical Research Laboratory, and Mr. John Kettlewell of Systems Research Laboratories, Inc., Dayton, Ohio. Mr. Furness served as project officer during development and application of the PIS simulator and Mr. Kettlewell was responsible for designing the electronic portion of the upgraded system. The research reported in this paper was documented by personnel of the Aerospace Medical Research Laboratory, Aerospace Medical Division, Air Force Systems Command, Wright-Patterson Air Force Base, Ohio. Reprints of this article are identified by Aerospace Medical Research Laboratory as AMRL-TR-78-60. Further reproduction is authorized to satisfy needs of the US Government.
PY - 2012/5
Y1 - 2012/5
N2 - We consider Markov decision processes where the values of the parameters are uncertain. This uncertainty is described by a sequence of nested sets (that is, each set contains the previous one), each of which corresponds to a probabilistic guarantee for a different confidence level. Consequently, a set of admissible probability distributions of the unknown parameters is specified. This formulation models the case where the decision maker is aware of and wants to exploit some (yet imprecise) a priori information of the distribution of parameters, and it arises naturally in practice where methods for estimating the confidence region of parameters abound. We propose a decision criterion based on distributional robustness: the optimal strategy maximizes the expected total reward under the most adversarial admissible probability distributions. We show that finding the optimal distributionally robust strategy can be reduced to the standard robust MDP where parameters are known to belong to a single uncertainty set; hence, it can be computed in polynomial time under mild technical conditions.
AB - We consider Markov decision processes where the values of the parameters are uncertain. This uncertainty is described by a sequence of nested sets (that is, each set contains the previous one), each of which corresponds to a probabilistic guarantee for a different confidence level. Consequently, a set of admissible probability distributions of the unknown parameters is specified. This formulation models the case where the decision maker is aware of and wants to exploit some (yet imprecise) a priori information of the distribution of parameters, and it arises naturally in practice where methods for estimating the confidence region of parameters abound. We propose a decision criterion based on distributional robustness: the optimal strategy maximizes the expected total reward under the most adversarial admissible probability distributions. We show that finding the optimal distributionally robust strategy can be reduced to the standard robust MDP where parameters are known to belong to a single uncertainty set; hence, it can be computed in polynomial time under mild technical conditions.
KW - Distributional robustness
KW - Markov decision process
KW - Parameter uncertainty
UR - https://www.scopus.com/pages/publications/84861386911
U2 - 10.1287/moor.1120.0540
DO - 10.1287/moor.1120.0540
M3 - Article
SN - 0364-765X
VL - 37
SP - 288
EP - 300
JO - Mathematics of Operations Research
JF - Mathematics of Operations Research
IS - 2
ER -