潜在博弈与乘积单纯形优化中的自界遗憾匹配+算法
Self-Bounding Regret Matching+ in Potential Games and Product-Simplex Optimization
浏览论文内容
中文总结 AI 辅助
本文提出RM+的精确一步守恒律,证明其在乘积单纯形光滑目标上可经O(ε⁻²)次迭代找到ε-KKT点,解决有限精确潜在博弈交替策略下遗憾有界的开放问题,性能优于对比变体。
中文摘要 AI 辅助
遗憾匹配+(RM+)无需参数、尺度不变,是大规模博弈求解的核心算法,但其仅有的个体遗憾通用保证随√T增长。ICLR近期一项研究利用该包络证明,RM+在乘积单纯形上的光滑目标函数可经O(ε⁻⁴)次迭代达到ε-平稳点,若采用标准零初始化则需O(ε⁻⁸)次迭代。本文给出RM+的精确一步守恒律,即正向效用收益同时支撑状态平方运动与遗憾-状态范数增长;对于m个动作,范数增长最多为正向收益的√(m-1)倍,且该系数是紧的。基于此,本文为未修改的RM+得出四项结果:1.其在任意效用路径上的遗憾受中心化时间变化控制;2.在所有有限精确潜在博弈的交替策略下,其遗憾被一致有界,解决了一个开放问题并使平方激活间隙可求和;3.经认证的惰性RM+与普通循环RM+均达到ε⁻²指数;4.在任意光滑、可能非凹的单纯形目标上,RM+可经O(ε⁻²)次迭代找到ε-KKT点。更广泛地,对于任意乘积单纯形上的光滑目标,循环块RM+从任意初始化出发也达到相同的O(ε⁻²)指数,且带有显式的轨迹依赖常数。证明过程控制了低状态块导致的有限目标损失,随后自界每个块状态及总平方路径长度,完整证明覆盖零状态、紧性、共同剖面平稳性及鲁棒收益优势。在图形潜在博弈与稠密非凸目标上,经oracle归一化的诊断工具对比了RM+与预测型及光滑 extra-gradient 变体的性能。
英文摘要
Regret matching+ (RM+) is parameter free, scale invariant, and central to large game solving, but its only general individual-regret guarantee grows as $\sqrt{T}$. A recent ICLR result used this envelope to prove that RM+ reaches an $ε$-stationary point of a smooth objective over a product of simplices in $O(ε^{-4})$ iterations, or $O(ε^{-8})$ from the standard zero initialization. We give an exact one-step conservation law for RM+. It states that forward utility gain pays for both squared state motion and growth of the regret-state norm. Norm growth is at most $\sqrt{m-1}$ times forward gain for $m$ actions, and the coefficient is sharp. This yields four results for unmodified RM+. Its regret on any utility path is controlled by centered temporal variation. Its regret is uniformly bounded under alternating play in every finite exact potential game, resolving an open question and making squared activation gaps summable. Both certified lazy and ordinary cyclic play attain an $ε^{-2}$ exponent. On any smooth, possibly nonconcave simplex objective, RM+ finds an $ε$-KKT point in $O(ε^{-2})$ iterations. Most broadly, for a smooth objective over an arbitrary product of simplices, cyclic block RM+ attains the same $O(ε^{-2})$ exponent from arbitrary initialization, with an explicit trajectory-dependent constant. The proof controls the finite objective loss caused by low-state blocks and then self-bounds every block state and the total squared path length. Complete proofs cover zero states, sharpness, common-profile stationarity, and robust gain dominance. Oracle-normalized diagnostics compare RM+ with predictive and smooth extra-gradient variants on graphical potential games and dense nonconvex objectives.
发表机构
- Johns Hopkins University(约翰斯·霍普金斯大学)
- Miramar College(米拉马尔学院)
机构由 AI 辅助整理,请以论文原文为准。