arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00504cs.GTcs.AIcs.LGcs.SYeess.SYmath.OC

折扣马尔可夫博弈中的独立强化学习

Independent Reinforcement Learning in Discounted Markov Games

Asrin Efe Yorulmaz, Ugur Aydin, Tamer Basar

首次发表
浏览论文内容

中文总结 AI 辅助

针对折扣一般和马尔可夫博弈,在假设ETH与PPAD等价下证明独立学习无多项式时间算法,提出首个完全解耦的分层乐观镜像下降算法,保证次指数收敛并开发全反馈与部分反馈版本。

中文摘要 AI 辅助

本研究探讨折扣一般和马尔可夫博弈中完全解耦的学习问题。假设“ETH 与 PPAD 等价”,我们证明,对于每个固定的折扣因子,当玩家在去中心化环境中独立学习时,不存在多项式时间算法能在折扣一般和马尔可夫博弈中计算逆多项式精度的粗相关均衡。作为该难度结果的补充,我们提出了首个完全解耦算法,该算法在不对博弈施加任何结构限制的情况下,保证了向粗相关均衡的次指数收敛。我们的算法是分层变体的乐观镜像下降,采用针对多智能体环境定制的递增步长策略。最后,我们开发了上述算法的全反馈和部分反馈版本,并分别为每种情况建立了次指数收敛保证。

英文摘要

In this work, we study radically uncoupled learning in discounted general-sum Markov games. Assuming ``$\mathsf{ETH}$ for $\mathsf{PPAD}$", we show that, for every fixed discount factor, there is no polynomial-time algorithm for computing inverse-polynomially accurate coarse correlated equilibria in discounted general-sum Markov games when players learn independently in decentralized settings. Complementing this hardness result, we provide what appears to be the first \emph{radically uncoupled} algorithm with sub-exponential convergence guarantees to coarse correlated equilibria in discounted general-sum Markov games without imposing any structural restrictions on the game. Our algorithm is a \emph{layered} variant of optimistic mirror descent with an increasing step-size schedule tailored to the multi-agent setting. Finally, we develop both full-feedback and partial feedback versions of the aforementioned algorithm and establish sub-exponential convergence guarantees for each case.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑