发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对未知独立链随机博弈,提出完全在线分散的镜像下降算法,以O(T^{-1/2})遗憾率逼近平稳NE,避免指数复杂度,并给出近似粗相关均衡保证。
AI 中文摘要
我们考虑具有独立受控链和未知转移核的随机博弈,其中玩家仅观察其局部状态和实现的收益。我们开发了一种完全在线、分散且无协调的镜像下降算法,该算法在占用测度的对偶空间中运行,以逼近平稳纳什均衡(NE)策略。该算法在每个原始时间步使用单个转移/奖励样本,仅依赖局部信息,并且既不需要联合状态空间的覆盖,也不需要同步的回合。在均匀遍历性和有限覆盖假设下,我们证明,以高概率,时间平均的固定比较器遗憾以规范的$O(T^{-1/2})$速率衰减,直至对数因子和游戏参数的幂次依赖。特别地,复杂度取决于各个局部状态空间的覆盖时间,而不是乘积状态空间,从而避免了对玩家数量和联合状态与动作空间大小的指数依赖。所得的有限时间遗憾界进一步产生了近似粗相关均衡保证,这对于任意奖励函数是自然的,因为在此设置中计算平稳$\epsilon$-NE是PPAD难的。在额外的全局变分稳定性条件下,我们证明相同的完全在线算法在最后一次迭代中渐近收敛到平稳$\epsilon$-NE。我们的结果为具有未知独立链的随机博弈提供了一个完全在线且可扩展的学习框架。该算法也可以被视为马尔可夫博弈的原始-对偶框架,该框架利用了玩家受控转移链的独立性和局部结构。
英文摘要
We consider stochastic games with independent controlled chains and unknown transition kernels, where players observe only their local states and realized payoffs. We develop a fully online, decentralized, and uncoordinated mirror-descent algorithm that operates in the dual space of occupancy measures for approximating stationary Nash equilibrium (NE) policies. The algorithm uses a single transition/reward sample at every primitive time step, relies only on local information, and requires neither coverage of the joint state space nor synchronized episodes. Under uniform-ergodicity and finite-coverage assumptions, we show that, with high probability, the time-averaged fixed-comparator regret decays at the canonical $O(T^{-1/2})$ rate, up to logarithmic factors and polynomial dependence on the game parameters. In particular, the complexity depends on the cover times of the individual local state spaces rather than the product state space, avoiding exponential dependence on the number of players and the sizes of the joint state and action spaces. The resulting finite-time regret bound further yields an approximate coarse-correlated-equilibrium guarantee, which is natural for arbitrary reward functions since computing a stationary $ε$-NE is PPAD-hard in this setting. Under an additional global variational-stability condition, we show that the same fully online algorithm converges asymptotically in the last iterate to a stationary $ε$-NE. Our results provide a fully online and scalable learning framework for stochastic games with unknown independent chains. The algorithm can also be viewed as a primal-dual framework for Markov games that exploits the independence and local structure of the players' controlled transition chains.