arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09780cs.GTmath.OC

零和博弈中的算法非对称性:针对缓慢对手的快速收敛的单侧恢复

Algorithmic Asymmetry in Zero-Sum Games: Unilateral Recovery of Fast Convergence Against a Slow Opponent

James P. Bailey, Soham Das

AI总结:

本文研究零和博弈的算法非对称性,提出智能体用交替乐观梯度下降(AOGD)可补偿普通梯度下降(GD)对手的缓慢收敛,使联合动态以O(1/T)速率收敛到纳什均衡。

AI中文摘要:

零和博弈中的学习动态通常在算法对称性下进行分析:两个智能体使用相同的更新规则,或来自同一算法族的方法。这与零和博弈的本质不符;竞争智能体无需在算法选择上协调一致。本文研究零和博弈学习动态中的算法非对称性,具体而言,当一个智能体固定使用普通梯度下降(vanilla gradient descent)时,其基于遗憾的标准分析最多仅能保证O(1/√T)的遍历收敛,我们探究是否可恢复快速收敛。我们证明缓慢的速率并非固有的:当一个智能体使用梯度下降时,对立智能体可使用一种改进的乐观更新方法,我们称之为交替乐观梯度下降(Alternating Optimistic Gradient Descent, AOGD),使联合动态在偶数迭代上模拟交替梯度下降。结果,梯度下降(GD)与AOGD的非对称动态的时间平均以O(1/T)的速率收敛到纳什均衡。我们的结果表明,快速收敛无需协调算法选择:一个智能体可对缓慢的对手进行补偿;更广泛地说,本文强调算法非对称性是理解多智能体优化中跨类交互的有用视角。

英文摘要:

Learning dynamics in zero-sum games are typically analyzed under algorithmic symmetry: both agents use the same update rule, or methods from a common algorithmic family. This is at odds with the nature of zero-sum games; competing agents need not coordinate on algorithm selection. This paper studies algorithmic asymmetry in learning dynamics in zero-sum games. In particular, we ask whether fast convergence can be recovered when one agent is fixed to vanilla gradient descent, whose standard regret-based analysis certifies, at best, $O(1/\sqrt{T})$ ergodic convergence. We show that the slow rate is not intrinsic. When one agent uses gradient descent, the opposing agent can use a modified optimistic update, which we call Alternating Optimistic Gradient Descent (AOGD), to make the joint dynamics simulate Alternating Gradient Descent on the even iterates. As a result, the time-average of the asymmetric GD vs.\ AOGD dynamics converges to Nash equilibria at rate $O(1/T)$. Our results show that fast convergence need not require coordinated algorithm selection: one agent can compensate for a slower opponent. More broadly, the paper highlights algorithmic asymmetry as a useful lens for understanding cross-class interactions in multiagent optimization.

↑