arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03621cs.GTcs.CCcs.DScs.LG

马尔可夫博弈中的正规型关联

Normal-Form Correlation in Markov Games

  • Carnegie Mellon University(卡内基梅隆大学)
  • Massachusetts Institute of Technology(麻省理工学院)
  • University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Strategy Robot, Inc.(Strategy Robot公司)
  • Strategic Machine, Inc.(Strategic Machine公司)
  • Optimized Markets, Inc.(Optimized Markets公司)

机构由 AI 辅助整理,请以论文原文为准。

Ioannis Anagnostides, Constantinos Daskalakis, Gabriele Farina, Noah Golowich, Tuomas Sandholm, Brian Hu Zhang

中文总结 AI 辅助

本文针对有限时域马尔可夫博弈,提出首个计算正规型关联均衡的高效算法,通过逆向归纳和常数期望关联均衡实现,并证明在多人或高精度情形下为PPAD完备。

中文摘要 AI 辅助

近年来,关于马尔可夫博弈中关联均衡概念的研究激增。然而,现有结果集中于比正规型关联均衡(NFCEs)更弱的概念,使得计算此类均衡这一更具挑战性的问题悬而未决,该问题可追溯至Papadimitriou和Roughgarden的开创性工作(JACM'08)。在此,我们为具有固定玩家数量$n$的有限时域马尔可夫博弈建立了首个高效的NFCEs算法。特别地,对于具有$S$个状态、时域$H$以及每位玩家至多$A$个动作的博弈,该算法在时间$S(AH/\epsilon)^{O(n)}$内计算出一个$\epsilon$-NFCE。这是首个在正规型设置之外的一类有趣问题中,关于NFCEs的关于$1/\epsilon$和博弈描述的多项式时间算法。此外,在推荐跨状态独立的通常假设下,我们证明了PPAD完备性——即与纳什均衡的计算等价性——要么在多人博弈中成立,要么在精度指数级小时成立。我们方法的关键思想是对一系列辅助阶段博弈进行逆向归纳,但不同之处在于每一步我们计算一个常数期望关联均衡。这是关联均衡的一种自然细化,其中服从建议的条件期望收益与建议本身无关。事实上,我们的归约是双向的,建立了马尔可夫博弈中常数期望CEs与NFCEs之间的等价性。对于固定数量的玩家,我们观察到常数期望CE可以通过结合线性规划和适当的离散化来近似计算。相比之下,在以下情况下它是PPAD难的:i)多项式矩阵(多人)博弈在常数精度下,以及ii)两人博弈在指数级小精度下。后一个结果源于与秩为2的两人博弈的意外联系。

英文摘要

There has been a surge of recent work on correlated equilibrium concepts in Markov games. However, existing results focus on concepts weaker than normal-form correlated equilibria (NFCEs), leaving open the more challenging question of computing such equilibria, which goes back to the seminal work of Papadimitriou and Roughgarden (JACM'08). Here, we establish the first efficient algorithm for NFCEs in finite-horizon Markov games with a fixed number of players $n$. In particular, with $S$ states, horizon $H$, and at most $A$ actions per player, it computes an $ε$-NFCE in time $S(AH/ε)^{O(n)}$. This is the first algorithm polynomial in $1/ε$ and the description of the game for NFCEs in an interesting class of problems beyond the normal-form setting. Moreover, under the usual assumption that recommendations are independent across states, we show PPAD-completeness---that is, computational equivalence to Nash equilibria---either in many-player games or when the precision is exponentially small. The key idea behind our approach is to run backward induction on a sequence of auxiliary stage games, but with the twist that in each step we compute a constant-expectation correlated equilibrium. This is a natural refinement of correlated equilibrium in which the conditional expected payoff from obeying is independent of the recommendation. In fact, our reduction goes both ways, establishing an equivalence between constant-expectation CEs and NFCEs in Markov games. For a fixed number of players, we observe that a constant-expectation CE can be computed approximately by combining linear programming with suitable discretization. In contrast, it is PPAD-hard in i) polymatrix (many-player) games at constant precision, and ii) two-player games at exponentially small precision. The latter result follows from an unexpected connection to rank-2 two-player games.

↑