马尔可夫链上的目标折扣和问题及其在马尔可夫决策过程中的应用
The Stochastic Target Discounted-Sum Problem
- Univ Rennes, Inria, CNRS, IRISA(雷恩大学、Inria、CNRS、IRISA)
- University of Sydney(悉尼大学)
- Université Libre de Bruxelles(布鲁塞尔自由大学)
- Rice University(莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究马尔可夫链上的目标折扣和问题,证明特定路径事件概率为零,用自动机理论求解,并将结果应用于马尔可夫决策过程,实现相关值的伪多项式时间计算与策略优化。
AI中文摘要:
折扣和是一种对有限字母表Σ上的权重序列进行聚合的方式,即对于折扣因子λ,序列w₀w₁w₂…∈Σ的折扣和为∑_{i∈ℕ}w_iλ^i。当前仍未解决的目标折扣和问题是:给定λ、Σ和目标值t,是否存在Σ上的无限序列,其折扣和等于t。我们研究并解决了该问题的概率变体,即马尔可夫链上的目标折扣和问题。为此,我们证明了包含折扣和等于目标值且具有无限多个不同后缀和的路径的事件概率为零。这一结构特性使我们能够使用自动机理论技术求解马尔可夫链上的目标折扣和问题。我们将技术结果应用于具有目标折扣和目标的马尔可夫决策过程:证明了下确界值和有限记忆上确界值可在伪多项式时间内计算,且由确定性有限记忆策略达到。
英文摘要:
The target discounted-sum problem (TDS) asks, given a finite integer alphabet $Σ$, a rational discount factor $λ$, and a rational target $t$, whether some infinite sequence over $Σ$ has discounted sum exactly $t$. This problem remains open and underlies several open questions in automata theory, games, and Markov decision processes. We introduce and solve its stochastic counterpart, the stochastic target discounted-sum problem, which replaces existence by computation of the probability. We show that the probability that a random sequence generated by a finite Markov chain has discounted sum $t$ is rational and computable in pseudo-polynomial time. We further show how to decide, in polynomial time, whether the discounted-sum distribution of a Markov chain is atomless, and how to approximate to an arbitrary precision the probability that the discounted sum exceeds a rational threshold. Our techniques for the stochastic TDS problem allow us to make progress on TDS objectives in stochastic games, which are known to be as hard as the TDS problem. Restricting the maximizing player to finite-memory strategies, while allowing the minimizing player to use arbitrary strategies, we reduce the value problem and the synthesis problem to corresponding problems for safety objectives in stochastic games. This yields computable optimal values and deterministic optimal strategies with pseudo-polynomially bounded memory for stochastic games, and results in pseudo-polynomial-time algorithms for special cases of Markov decision processes and deterministic two-player games.