AI 中文总结
本文证明在一般和博弈中,乐观 Hedge 算法采用常数步长可在自我对弈中实现 O(√n log d_i log T) 的对数级个体遗憾,优于对抗性遗憾界,并通过高阶差分分析改进现有结果。
AI 中文摘要
简单的无遗憾动态在自我对弈中能否比对抗任意对手获得更小的遗憾?在 n 人一般和博弈中,Daskalakis 等人 2021 证明了乐观 Hedge 算法的 O(n log d_i log^4 T) 个体遗憾界,这改进了经典的 O(√T) 对抗遗憾界。在本工作中,我们证明在期望损失向量反馈下,采用常数步长的乐观 Hedge 算法可以进一步实现 O(√n log d_i log T) 的个体外部遗憾。相应地,时间平均策略的粗相关均衡间隙为 O(√n log d log T / T),其中 d = max_i d_i。这一改进源于更大的可容许步长 η = Θ(1/(√n log T))。我们的分析证明了概率加权成对损失间隙的高阶差分的阶乘界,然后在固定欧几里得范数下应用有限差分插值。这些估计强化了 Daskalakis 等人 2021 的分析,并得出了对数遗憾界。
英文摘要
Can simple no-regret dynamics attain smaller regret in self-play than against arbitrary adversaries? In $n$-player general-sum games, Daskalakis et al. 2021 proved an $O(n\log d_i\log^4 T)$ individual regret bound for Optimistic Hedge, which improves upon the classical $O(\sqrt T)$ adversarial regret bound. In this work, we show that Optimistic Hedge with a constant step size can further achieve $O(\sqrt n\log d_i\log T)$ individual external regret under expected loss-vector feedback. The time-averaged play consequently enjoys a coarse correlated equilibrium gap $O(\sqrt n\log d\log T/T)$, where $d=\max_i d_i$. The improvement comes from a larger admissible step size $η=Θ(1/(\sqrt n\log T))$. Our analysis proves factorial bounds on high-order differences of probability-weighted pairwise loss gaps, then applies finite-difference interpolation in a fixed Euclidean norm. These estimates sharpen the analysis of Daskalakis et al. 2021 and yield a logarithmic regret bound.