Parameterized Markov decision process and its application to service rate control

Li Xia ^[1] ; Qing-Shan Jia ^[1]
1. [1] Tsinghua University
  
  Tsinghua University
  
  China
Localización: Automatica: A journal of IFAC the International Federation of Automatic Control, ISSN 0005-1098, Vol. 54, 2015, págs. 29-35
Idioma: inglés
Texto completo no disponible (Saber más ...)
Resumen
- In this paper, we discuss the optimization of Markov decision processes (MDPs) with parameterized policy, where the state space is partitioned and a parameter is assigned to each partition. The goal is to find the optimal parameters which maximize the long-run average performance. The traditional policy iteration is usually inapplicable to parameterized policy because the parameter tuning at different states are correlated. With some appropriate assumptions and special conditions, we develop a modified policy iteration type algorithm to find the optimal parameters. Compared with the traditional gradient-based approaches for MDP with parameterized policy, this policy iteration type approach is much more efficient. Finally, as an example, we apply this approach to a service rate control problem in closed Jackson networks. As compared with the gradient-based approach which is trapped into local optimum, our approach is demonstrated to efficiently find the optimal service rates in global scope.