多智能体强化学习中的均衡
Equilibrium in Multi-Agent Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本文针对多智能体强化学习,提出马尔可夫贝叶斯粗相关均衡(MBCCE),证明自适应马尔可夫粗悔憾(AMCR)消失时经验分布聚点为MBCCE,且两种算法可生成近似MBCCE并得到有限时间收敛速率。
中文摘要 AI 辅助
随机博弈的标准解概念,如马尔可夫完美均衡和马尔可夫粗相关均衡,在计算上十分困难,因此通常不能期望标准的去中心化强化学习算法会收敛到这些均衡。本文研究此类算法生成的均衡。特别地,我们为随机博弈引入一种新的解概念:马尔可夫贝叶斯粗相关均衡(Markov Bayes coarse correlated equilibrium, MBCCE),其定义为状态与平稳策略剖面的分布,满足在观察到状态后但在观察到推荐的动作前,没有玩家通过选择不同的当前动作能获益,且后续博弈由采样得到的策略剖面支配。我们讨论MBCCE与有限标准型博弈中的粗相关均衡(coarse correlated equilibrium, CCE)之间的相似性,并证明MBCCE保留了CCE的若干关键性质。随后我们引入对应的悔憾概念:自适应马尔可夫粗悔憾(adaptive Markov coarse regret, AMCR),并证明AMCR消失意味着已实现状态与策略剖面的经验分布的每个聚点都是MBCCE。关键的是,我们证明实现AMCR可简化为两个标准学习任务:在每个状态下最小化外部悔憾,以及准确评估当前联合策略。接着我们证明在温和条件下,这些性质对两种自然的强化学习算法设计成立:一种去中心化异步演员-评论家算法,通过新的双时间尺度随机近似分析;以及一种标准的情节式多智能体投影策略梯度方法。因此,两种算法均能生成近似MBCCE,且我们为两者建立了明确的有限时间收敛速率。
英文摘要
Standard solution concepts for stochastic games, such as Markov perfect equilibrium and Markov coarse correlated equilibrium, are computationally difficult, and thus, standard decentralized reinforcement-learning algorithms should not generally be expected to converge to them. In this paper, we study the equilibrium generated by such algorithms. In particular, we introduce a new solution concept for stochastic games, Markov Bayes coarse correlated equilibrium (MBCCE), defined as a distribution over states and stationary policy profiles such that, after observing the state but before observing her recommended action, no player can gain by choosing a different current action, with the sampled policy profile governing play thereafter. We discuss the parallels between MBCCE and coarse correlated equilibrium (CCE) in finite normal-form games and show that MBCCE retains several of its key properties. We then introduce a corresponding regret notion, adaptive Markov coarse regret (AMCR), and show that vanishing AMCR implies that every accumulation point of the empirical distribution of realized states and policy profiles is an MBCCE. Crucially, we show that achieving AMCR reduces to two standard learning tasks: minimizing external regret at each state and accurately evaluating the current joint policy. We then prove that under mild conditions these properties hold for two natural RL algorithmic designs: a decentralized asynchronous actor--critic algorithm through a new two-timescale stochastic-approximation analysis, and a standard episodic multi-agent projected policy-gradient method. Hence, both algorithms generate approximate MBCCEs, and we establish explicit finite-time convergence rates for both.