arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24986cs.AI

通过部分可观察随机博弈求解具有ω-正则目标的鲁棒部分可观察马尔可夫决策过程

Solving Robust POMDPs with Omega-regular Objectives via Partially Observable Stochastic Games

  • Indian Institute of Technology Bombay(印度孟买理工学院)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Durgam Latha, Dion Reji, S. Akshay, Đorđe Žikelić, Shankaranarayanan Krishna

AI总结:

本研究建立了具有多面体不确定集合的(s,a)-矩形鲁棒部分可观察马尔可夫决策过程与部分可观察随机博弈的语义等价性,推导了求解具有ω-正则目标的RPOMDPs及RMDPs的新计算复杂度结果。

AI中文摘要:

鲁棒部分可观察马尔可夫决策过程(RPOMDPs)将经典POMDPs推广到精确转移概率未知的场景,仅已知其属于某个值的不确定集合。本研究致力于求解具有通用ω-正则目标的RPOMDPs,这类目标涵盖可达性、安全性及线性时序逻辑(LTL)目标等广泛类别。我们证明,对于具有多面体不确定集合的(s,a)-矩形RPOMDPs,求解其ω-正则目标的问题可归约为求解具有ω-正则目标的部分可观察随机博弈(POSGs)。此外,我们首次证明可在两个方向构建归约关系,建立了具有多面体不确定集合的(s,a)-矩形RPOMDPs与POSGs之间的语义等价性。这使我们能够推导得出一系列新的计算复杂度结果,包括求解具有不同ω-正则目标的RPOMDPs的上下复杂度界。作为推论,我们还推导了鲁棒马尔可夫决策过程(RMDPs)的新计算复杂度结果。

英文摘要:

Robust POMDPs (RPOMDPs) generalize classical POMDPs to the setting where exact transition probabilities are not known -- rather, they are only known to belong to some uncertainty set of values. In this work, we study the problem of solving RPOMDPs with general omega-regular objectives, which subsume a broad class of objectives such as reachability, safety, and linear temporal logic (LTL) objectives. We show that, for (s,a)-rectangular RPOMDPs with polytopic uncertainty sets, the problem of solving RPOMDPs under omega-regular objectives can be reduced to solving partially observable stochastic games (POSGs) under omega-regular objectives. Moreover, we show for the first time that reductions can be constructed in both directions, establishing the semantic equivalence between (s,a)-rectangular RPOMDPs with polytopic uncertainty sets and POSGs. This allows us to derive a range of new computational complexity results, including both upper and lower complexity bounds, on solving RPOMDPs with different omega-regular objectives. As a corollary, we also derive new computational complexity results for RMDPs.

↑