arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

零和网络化可分离马尔可夫博弈中策略动力学的末次迭代收敛

Last-Iterate Convergence of Policy Dynamics in Zero-Sum Networked Separable Markov Games

Zailin Ma

arXiv 2609.08823首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对零和网络化可分离马尔可夫博弈,提出熵正则化乐观乘法权重更新算法,实现末次迭代近似纳什均衡,并证明近线性收敛速率。

AI 中文摘要

求解一般多人马尔可夫博弈的纳什均衡在计算上是棘手的,而两人零和马尔可夫博弈则允许快速的末次迭代策略优化方法。有限时域零和网络化可分离马尔可夫博弈占据了一个重要的中间地带:它们通过成对交互保留了全局竞争结构,同时在完全信息和已知转移设置下保持了纳什均衡(NE)的计算可处理性。针对此类博弈的现有算法要么通过均衡坍缩论证来处理一个简化设置(其中单一控制器决定转移概率),要么依赖每阶段均衡求解器的反向动态规划。然而,直接策略更新方法的设计与分析仍然不足。为解决此问题,我们提出了熵正则化乐观乘法权重更新(ER-OMWU),这是一种互补的单循环策略动力学,它对称地更新玩家的策略,并在末次迭代中返回近似纳什均衡。我们首次提供了所关注博弈中策略动力学的末次迭代收敛分析:经过$\widetilde{O}(1/{\epsilon})$次迭代后,返回的策略是一个$\epsilon$-近似纳什均衡。该结果保留了两人零和马尔可夫博弈中策略优化所达到的近线性收敛速率,但将策略动力学视角扩展到了更复杂但结构化的多人设置。

英文摘要

Solving Nash equilibria for general multi-player Markov games is computationally intractable, while two-player zero-sum Markov games admit fast last-iterate policy-optimization methods. Zero-sum networked separable Markov games occupy an important middle ground: they retain global multi-player competition structure through pairwise interactions, while preserving computational tractability of Nash equilibria (NE) in the finite-horizon setting. Existing algorithms for this class either proceed through equilibrium-collapse arguments for a simplified setting where a single controller determines the transition probability, or backward dynamic programming that relies on equilibrium solvers at each stage. However, the design and analysis of direct policy-update approaches remain inadequate. To address this issue, we propose the entropy-regularized optimistic multiplicative weights update (ER-OMWU), a complementary single-loop policy dynamic that updates players' policies symmetrically and returns an approximate NE in the last iteration. We provide the first last-iterate convergence analysis of policy dynamics in the games of interest: after $\widetilde O\left(1/ε\right)$ iterations, the returned policy is an $ε$-approximate Nash equilibrium. The result preserves the near-linear convergence rate achieved by policy optimization in two-player zero-sum Markov games, but extends the policy-dynamics viewpoint to a more complicated but structured multi-player setting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑