带长期平均目标的可观测部分马尔可夫决策过程(POMDP)中值的近似复杂度
The Complexity of Approximating the Value in Revealing POMDPs with Long-Run Average Objectives
浏览论文内容
中文总结 AI 辅助
该研究针对带长期平均目标的可观测POMDP,明确其长期平均值近似问题为EXPTIME完全问题,还通过控制优化实例说明该类POMDP的实际应用价值。
中文摘要 AI 辅助
我们研究具有长期平均目标的部分可观测马尔可夫决策过程(POMDP),其定义为期望平均奖励的下极限。一般而言,POMDP的长期平均值既不可计算也不可近似。因此,我们考虑一类特殊的“可观测POMDP”,在该类中,每一步都有正概率将当前状态反馈给控制器。首先,我们通过控制与优化领域的一个应用实例说明该类的实际意义;其次,我们证明,近似具有长期平均目标的可观测POMDP的长期平均值是EXPTIME完全问题,从而确定了其严格的计算复杂度。
英文摘要
We study partially observable Markov decision processes (POMDPs) with long-run average objectives, where the payoff is defined as the limit inferior of the expected average rewards. We consider the computational problem of approximating the long-run average value of a POMDP. In general, the long-run average value of a POMDP is neither computable nor approximable. We therefore consider the subclass of revealing POMDPs. Informally, a POMDP is revealing when the controller observes the underlying state with positive probability at each stage of the process. Our main contributions are threefold. First, we illustrate the practical relevance of this class of POMDPs through an application in control and optimization. Second, we present an exponential-time algorithm for the value approximation problem. Third, we establish EXPTIME-hardness by a reduction from the problem of almost-sure safety in POMDPs. Together, these results show that the problem of approximating the long-run average value in revealing POMDPs is EXPTIME-complete.
发表机构
- Institute of Science and Technology Austria(奥地利科学技术学院)
机构由 AI 辅助整理,请以论文原文为准。