发表机构
The University of Tokyo; RIKEN(东京大学; 理化学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多人一般和博弈,提出结合Blum–Mansour归约与混合正则化的无耦合学习动力学,实现次对数级交换遗憾,得到首个此类个体保证及近似相关均衡。
AI 中文摘要
交换遗憾控制着多人一般和博弈中无耦合学习动力学收敛到相关均衡的速率。在全信息反馈下,当每个玩家遵循相同动力学时,之前的最佳保证随时间范围T呈对数增长。我们构建了无耦合动力学,使得每个玩家仅产生O(nm²√(log m log T))的交换遗憾,其中n为玩家数量,m为每个玩家的动作数量上限。据我们所知,这是该设置下首个次对数级个体保证,意味着博弈的时间平均乘积分布是O(nm²√(log m log T)/T)近似相关均衡。关键算法选择是将Blum–Mansour归约与乐观正则化跟随领导者结合,使用混合正则化分别对负香农熵和对数障碍加权:熵控制乐观预测误差,对数障碍通过其Bregman散度控制转移矩阵的移动。针对马尔可夫链平稳分布的新敏感性定理(不涉及混合参数或最小转移概率)将此控制传递给博弈策略,产生更简单的分析,无需局部范数或自和谐论证。该保证可通过对抗鲁棒变体保留,该变体额外确保针对任意效用序列的O(nm²√(log m log T)+√(mT log m))交换遗憾,也可通过无需T先验知识的无时间范围变体保留。
英文摘要
Swap regret governs the rate at which uncoupled learning dynamics converge to correlated equilibria in multiplayer general-sum games. Under full-information feedback, the best previous guarantee when every player follows the same dynamics grows logarithmically in the horizon $T$. We construct uncoupled dynamics under which every player incurs only $O(nm^2\sqrt{\log m\log T})$ swap regret, where $n$ is the number of players and $m$ bounds the number of actions per player. To our knowledge, this is the first sublogarithmic individual guarantee in this setting, and it implies that the time-averaged product distribution of play is an $O(nm^2\sqrt{\log m\log T}/T)$-approximate correlated equilibrium. The key algorithmic choice is to combine the Blum--Mansour reduction with optimistic follow-the-regularized-leader using a hybrid regularizer that separately weights negative Shannon entropy and the log-barrier: the entropy controls the optimistic prediction error, whereas the log-barrier controls the transition-matrix movement through its Bregman divergence. A new sensitivity theorem for stationary distributions of Markov chains, which involves neither mixing parameters nor the smallest transition probability, transfers this control to the played strategies and yields a simpler analysis without local-norm or self-concordance arguments. The guarantee is preserved by an adversarially robust variant that additionally ensures $O(nm^2\sqrt{\log m\log T}+\sqrt{mT\log m})$ swap regret against arbitrary utility sequences, and by a horizon-free variant that requires no prior knowledge of $T$.
Comments25 pages