发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对马尔可夫α-势博弈,提出KL投影自然策略梯度算法,在情节式和完全在线异步设置下实现高概率纳什遗憾界,消除分布不匹配系数,并应用于马尔可夫拥塞博弈的在线作业调度。
AI 中文摘要
我们研究在强盗反馈下无限时域折扣马尔可夫博弈中纳什均衡(NE)的分散式学习,重点关注马尔可夫 $\alpha$-势博弈。我们开发了KL投影自然策略梯度(NPG)算法,适用于两种设置:一种是在采样期间策略冻结的情节式设置,另一种是完全在线设置,其中每个玩家在每个时间步接收单个实现的成本样本,并沿着持续轨迹异步更新其策略。我们分别建立了情节式和完全在线设置下有限时间高概率NE遗憾界,阶数为 $\widetilde O(T^{-1/4})$ 和 $\widetilde O(T^{-2/15})$,直至固定的近似项。关键的是,我们的界消除了分布不匹配系数(该系数可能随状态空间大小而急剧增长),同时容纳了势近似误差、估计预言机偏差和转移敏感性。我们进一步识别了一种状态相关的势结构,该结构产生了更优的保证,对势近似误差 $\alpha$ 具有加性依赖。我们将该框架专门应用于独立资源马尔可夫拥塞博弈(IMCGs),建立了它们的近似势和转移敏感性性质,并从实现的成本中构建了分散式估计预言机。作为应用,我们引入了随机机器上的战略在线作业调度,并获得了一种可扩展的分散式算法来学习稳定的调度策略。总体而言,我们的结果为马尔可夫 $\alpha$-势博弈中的完全在线异步分散式学习提供了首个有限时间高概率NE遗憾保证,从遗憾界中消除了分布不匹配系数,容纳了固定的估计预言机偏差,并为IMCGs提供了具有有限时间保证的可扩展分散式学习。
英文摘要
We study decentralized learning of Nash equilibria (NE) in infinite-horizon discounted Markov games under bandit feedback, focusing on Markov $α$-potential games. We develop KL-projected natural policy gradient (NPG) algorithms in two settings: an episodic setting with frozen policies during sampling and a fully online setting in which players receive a single realized cost sample per time step and update their policies asynchronously along a continuing trajectory. We establish finite-time high-probability NE regret bounds of order $\widetilde O(T^{-1/4})$ and $\widetilde O(T^{-2/15})$ for the episodic and fully online settings, respectively, up to fixed approximation terms. Crucially, our bounds eliminate the distribution-mismatch coefficient, which can scale prohibitively with the size of the state space, while accommodating potential approximation, estimation-oracle bias, and transition sensitivity. We further identify a state-wise potential structure that yields sharper guarantees with additive dependence on the potential approximation error $α$. We specialize the framework to independent-resource Markov congestion games (IMCGs), establish their approximate-potential and transition-sensitivity properties, and construct decentralized estimation oracles from realized costs. As an application, we introduce strategic online job scheduling on stochastic machines and obtain a scalable decentralized algorithm for learning stable dispatching policies. Overall, our results provide the first finite-time high-probability NE regret guarantees for fully online asynchronous decentralized learning in Markov $α$-potential games, remove distribution-mismatch coefficients from the regret bounds, accommodate fixed estimation-oracle bias, and provide scalable decentralized learning with finite-time guarantees for IMCGs.