发表机构
Institute of Science and Technology Austria(奥地利科学技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对线性定义不确定性集的(s,a)-矩形鲁棒马尔可夫决策过程,解决了ω-正则奇偶性目标的定量分析问题,证明存在纯无记忆最优策略并提出多项式时间算法,通过实验与随机博弈方法对比验证了效果。
AI 中文摘要
鲁棒马尔可夫决策过程(RMDP)通过允许转移概率存在不确定性并针对其最坏情况实现进行优化,对经典马尔可夫决策过程(MDP)进行了泛化。我们考虑具有线性定义不确定性集的(s,a)-矩形RMDP,并研究奇偶性目标,这是ω-正则目标的规范表示。若不确定性集由转移分布及辅助变量上的线性不等式描述,则称其为线性定义的,这类不确定性集涵盖标准的L₁球、L_∞球以及一般的多面体不确定性集。定量值是智能体所有策略中,针对对抗性环境保证满足目标的概率的上确界。过往研究关注定性分析,即询问是否存在单个智能体策略,能以概率1(或正概率)针对所有环境策略保证目标满足的几乎必然(或正)问题。本研究中我们解决了精确定量问题,贡献有三:其一,证明智能体和环境均存在纯无记忆最优策略;其二,针对线性定义的鲁棒马尔可夫链的定量奇偶性问题,给出多项式时间算法,并将其作为RMDP策略迭代算法的子例程,该算法结合了定量单步改进与定性几乎必然改进;其三,报告了将我们的方法与显式归约为随机博弈的对比实验结果。
英文摘要
Robust Markov Decision Processes (RMDPs) generalize classical MDPs by allowing uncertainty in transition probabilities and optimizing against their worst-case realization. We consider $(s,a)$-rectangular RMDPs with \emph{linearly defined} uncertainty sets and study parity objectives, which are a canonical representation of $ω$-regular objectives. An uncertainty set is linearly defined if it is described by linear inequalities over the transition distribution together with auxiliary variables, which capture the standard $L_1$ and $L_\infty$ balls as well as general polytopic uncertainty sets. The quantitative value is the supremum, over all agent policies, of the satisfaction probability guaranteed against the adversarial environment. Previous work studied the qualitative analysis, namely the almost-sure (resp. positive) problem that asks whether a single agent policy guarantees satisfaction with probability one (resp. positive probability) against every environment policy. In this work, we solve the exact quantitative problem. Our contributions are threefold. First, we show that both the agent and the environment admit pure memoryless optimal policies. Second, we give a polynomial-time algorithm for quantitative parity on linearly defined robust Markov chains and use it as a subroutine in a policy-iteration algorithm for RMDPs. The algorithm combines quantitative one-step improvements with qualitative almost-sure improvements. Finally, we report experiments comparing our approach with the explicit reduction to stochastic games.
Comments26 Pages