arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

鲁棒马尔可夫决策过程的线性规划表示与强多项式算法

Linear Programming Representations and Strongly Polynomial Algorithms for Robust Markov Decision Processes

Han Zhong, Yinyu Ye

arXiv 2610.02131首次发表:更新:

发表机构

Shanghai Jiao Tong University; SIMIS; Stanford University(上海交通大学; 上海数学与交叉学科研究院; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出鲁棒马尔可夫决策过程的线性规划表示,通过编码鲁棒策略迭代步骤构造单一LP,并给出强多项式算法及改进的复杂度界。

AI 中文摘要

我们研究了具有有理多面体状态-动作矩形不确定性的奖励和转移的鲁棒马尔可夫决策过程(RMDPs)的线性规划(LP)表示与强多项式算法。通过编码有限序列的鲁棒策略迭代步骤,我们构造了一个单一的LP,其最优解能够恢复鲁棒最优值和所有最优平稳随机化策略。在固定折扣下,该LP具有多项式维度和编码长度,并且可以在强多项式时间内构造。我们还开发了一个鲁棒策略迭代的一般复杂度分析,该分析将不确定性集上的最小化成本与评估策略所需的迭代次数相结合。对于固定的折扣因子,我们利用此分析改进了已知的$\ell_1$和$\ell_\infty$ RMDPs的复杂度界,并为一般区间、加权$\ell_1$和Wasserstein RMDPs以及具有这些不确定性集的回合制随机博弈建立了新的强多项式界。

英文摘要

We study linear programming (LP) representations and strongly polynomial algorithms for robust Markov decision processes (RMDPs) with rational polyhedral state-action rectangular uncertainty in rewards and transitions. By encoding a finite sequence of robust policy-iteration steps, we construct a single LP whose optimal solutions recover the robust optimal value and all optimal stationary randomized policies. At fixed discount, the LP has polynomial dimension and encoding length and can be constructed in strongly polynomial time. We also develop a general complexity analysis of robust policy iteration that combines the cost of minimizing over uncertainty sets with the number of iterations needed to evaluate a policy. For a fixed discount factor, we use this analysis to improve the known complexity bounds for $\ell_1$ and $\ell_\infty$ RMDPs and establish new strongly polynomial bounds for general interval, weighted $\ell_1$, and Wasserstein RMDPs, as well as turn-based stochastic games with these uncertainty sets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑