AI 中文总结
本文证明常步长乐观 Hedge 在一般和博弈中达到常数个体遗憾,并导出粗相关均衡间隙,通过实解析递推消除时间范围依赖。
AI 中文摘要
简单的学习规则能否在自对弈中保持其遗憾有界?近期工作通过修改正则化和高阶预测实现了常数遗憾界。然而,对于乐观 Hedge——博弈中最具典范性的方法——已知的最佳个体遗憾界仍是对数阶的。在本工作中,我们证明,在期望损失向量反馈下,具有常数步长的朴素乐观 Hedge 在具有 n 个玩家和 d=(d_1,…,d_n) 个动作的一般和博弈中可以达到 O_{n,d}(1) 的个体遗憾。作为推论,其时间平均玩法享有 O_{n,d}(1/T) 的粗相关均衡(CCE)间隙。我们的分析将乐观 Hedge 表示为紧空间上的实解析递推,这产生了精确的有限阶差分关系,从而消除了对时间范围的依赖。我们的证明依赖于 Frisch (1967) 的非构造性 Noetherianity 论证,因此 (n,d) 依赖性仍然是隐式的。
英文摘要
Can simple learning rules keep their regret bounded in self-play? Recent work achieves constant regret bounds through modified regularization and higher-order prediction. Yet for Optimistic Hedge, arguably the most canonical method in games, the best known individual regret bound remains logarithmic. In this work, we prove that plain Optimistic Hedge with a constant step size can attain $O_{n,d}(1)$ individual regret in general-sum games with $n$ players and $d=(d_1,\ldots,d_n)$ actions, under expected loss-vector feedback. As a corollary, its time-averaged play enjoys an $O_{n,d}(1/T)$ coarse correlated equilibrium (CCE) gap. Our analysis represents Optimistic Hedge as a real-analytic recurrence on a compact space, which yields an exact finite-order difference relation that eliminates horizon dependence. Our proof hinges on nonconstructive Noetherianity argument of Frisch (1967), so the $(n,d)$-dependence remains implicit.