发表机构
Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文揭示了IREG-PRM+在零和矩阵博弈中最后迭代线性收敛的机制——范数饱和,并给出了显式收敛速率条件,通过比率证书验证了其有效性。
AI 中文摘要
IREG-PRM+ 通过其自身的范数对累积遗憾向量进行归一化,并在无需知道收益尺度的情况下获得最优遗憾。在零和矩阵博弈上未经修改地运行时,它在最后迭代中线性收敛,而目前没有分析解释其原因。障碍在于该算法没有固定的步长可供分析:步长是一个状态变量,是轨迹本身移动的遗憾范数的倒数。所有已证明的遗憾匹配动力学的线性速率都来自重启或修改更新。我们将其机制识别为范数饱和:遗憾范数上升到有限极限并冻结步长。我们证明它总是如此,并给出显式界,且饱和迫使最后迭代的纳什间隙在每个矩阵博弈上消失;当均衡唯一时,逐点收敛随之而来。在唯一的严格互补均衡附近,活跃支撑集在一步内冻结,该支撑集上的单轮雅可比矩阵具有闭式形式。然后,最后迭代以闭式速率线性收敛,前提是一个尺度不变的量保持低于1:饱和步长乘以支撑集上以价值为中心的收益子矩阵的最大奇异值。在包含216个实例的测试平台上,184个具有可解析极限的实例均满足该条件。同样的分析给出了一个比率证书:可观测的范数增量比率以常数为界约束不可观测的纳什间隙比率,该常数只进入一次且不随迭代次数累积。其斜率二定律在斜率可测量的实例中占96.1%成立,并且相同的增量监测扩展式博弈中的进展,其中最佳响应传递可以稀疏调度。代码可在以下网址获取:此 https URL。
英文摘要
IREG-PRM+ normalizes the cumulative regret vector by its own norm and attains optimal regret without knowledge of the payoff scale. Run unmodified on zero-sum matrix games, it converges linearly in the last iterate, and no analysis explains why. The obstacle is that the algorithm has no fixed step size to analyze: the step size is a state variable, the inverse of a regret norm that the trajectory itself moves. Every proved linear rate for regret-matching dynamics comes from restarting or modifying the update. We identify the mechanism as norm saturation: the regret norm rises to a finite limit and freezes the step size. We prove that it always does, with an explicit bound, and that saturation forces the last-iterate Nash gap to vanish on every matrix game; pointwise convergence follows whenever the equilibrium is unique. Near a unique strictly complementary equilibrium the active support freezes in one step, and the one-round Jacobian on that support has a closed form. The last-iterate then converges linearly at a closed-form rate, provided one scale-invariant quantity stays below one: the saturated step size times the largest singular value of the value-centered payoff submatrix on the support. On the $216$-instance testbed, the $184$ instances with a resolvable limit all satisfy it. The same analysis gives a ratio certificate: observable norm-increment ratios bound the unobservable Nash-gap ratio up to a constant that enters once and does not accumulate with the iteration count. Its slope-two law holds on $96.1\%$ of the instances where the slope is measurable, and the same increment monitors progress in extensive-form games, where best-response passes can be scheduled sparsely. The code is available at https://github.com/lbn187/NormCert.
Comments50 pages, 2 figures