通过高阶乐观策略实现一般博弈中的常数遗憾
Constant regret in general games via higher-order optimism
浏览论文内容
中文总结 AI 辅助
本文提出带折扣的高阶乐观算法(HOOD),作为OptFTRL的变体,可在N玩家标准型博弈中保证O(N³log²K)的个体遗憾,消除了此前实现一般博弈常数遗憾的关键障碍,且与同期独立工作有相似性。
中文摘要 AI 辅助
我们提出一种非耦合学习算法,当任意N玩家标准型博弈的所有玩家均采用该算法(每个玩家最多有K个行动)时,可保证在整个博弈过程中个体遗憾为O(N³log²K)。所提出的算法称为带折扣的高阶乐观算法(HOOD),是乐观正则化跟随领导者算法(OptFTRL)的变体,它将折扣后的(N+1)阶预测器与博弈策略空间合适“提升”上的熵正则化相结合。这种组合设计旨在以可控方式抑制诱导博弈序列的大幅振荡,从而消除了此前尝试在一般博弈中实现常数遗憾的一个关键障碍。我们的方法与同期完全独立的Liu、Farina和Ozdaglar的工作(arXiv:2608.31166)存在显著相似之处,他们近期通过使用高阶乐观策略和指数移动平均估计器,推导出了O(N²¹log⁴K)的遗憾界。
英文摘要
We introduce an uncoupled learning algorithm which, when employed by all players of an arbitrary $N$-player normal form game with up to $K$ actions per player, guarantees $O(N^3\log^2 K)$ individual regret, uniformly over the horizon of play. The proposed algorithm - which we call higher-order optimism with discounting (HOOD) is a variant of optimistic follow-the-regularized-leader (OptFTRL) that combines a discounted $(N+1)$-th order predictor with entropic regularization over a suitable "lifting" of the game's strategy space. This combination of ingredients is purposefully designed to dampen large oscillations of the induced sequence of play in a controlled manner, removing in this way a key stumbling block of previous attempts to achieve constant regret in general games. Our approach bears several striking similarities to the concurrent - and completely independent - work of Liu, Farina, and Ozdaglar (arXiv:2608.31166), who very recently derived an $O(N^{21}\log^{4} K)$ regret bound through the use of higher-order optimism and an exponential moving average estimator.
发表机构
- Univ. Mohammed VI Polytechnic(穆罕默德六世大学)
- Univ. Grenoble Alpes(格勒诺布尔大学)
- CNRS(法国国家科学研究中心)
- Inria(法国国家信息与自动化研究所)
- Grenoble INP(格勒诺布尔工业大学)
- LIG(信息与信号处理实验室)
机构由 AI 辅助整理,请以论文原文为准。