发表机构
MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出精确性作为统一原则,几何刻画了梯度下降、镜像下降和FTRL无遗憾的偏差类别,并引入保守相关均衡,揭示不同一阶方法控制不同偏差类。
AI 中文摘要
在线梯度下降通常通过外部遗憾来研究,其中学习器与固定的备选方案竞争。近期工作表明,一阶方法控制着更丰富的依赖于行动的偏差。我们寻求对偏差的几何刻画,使得在线梯度下降、镜像下降和跟随正则化领导者(FTRL)相对于这些偏差实现无遗憾。我们将精确性识别为共同原则。精确性意味着相关的位移场由标量势生成,或者等价地,相关的微分形式在算法所使用的几何中是精确的。该几何取决于算法。对于梯度下降,它是欧几里得几何;对于镜像下降,它是由正则化器诱导的几何;对于FTRL,它是累积对偶状态。在温和的正则性条件下,精确性产生次线性遗憾,而非零环量提供互补的障碍并导致线性遗憾。这为理解这些算法所控制的偏差类别提供了一个统一的几何框架,并揭示不同的一阶方法可以控制真正不同类别的偏差。这些偏差类别对学习有直接影响,特别是在博弈中。我们研究了由精确形式偏差诱导的均衡概念,并引入了保守相关均衡,反映了底层位移场的保守几何以及玩家可用的受限偏差族。我们刻画了其与相关均衡的关系,确定了所得均衡概念何时重合、何时分离,并展示了这些关系如何依赖于几何和学习算法。总体而言,这项工作为第一阶在线学习算法相对于固定比较器之外的内容实现无遗憾提供了一个统一的几何解释。
英文摘要
Online gradient descent is usually studied through external regret, where the learner competes with fixed alternatives. Recent work shows that first-order methods control richer action-dependent deviations. We ask for a geometric characterization of the deviations with respect to which online gradient descent, mirror descent, and follow-the-regularized-leader (FTRL) achieve no regret. We identify exactness as the common principle. Exactness means that the relevant displacement field is generated by a scalar potential, or equivalently that the associated one-form is exact in the geometry used by the algorithm. This geometry depends on the algorithm. For gradient descent it is Euclidean geometry, for mirror descent it is the geometry induced by the regularizer, and for FTRL it is the cumulative dual state. Under mild regularity conditions, exactness yields sublinear regret, while nonzero circulation provides the complementary obstruction and leads to linear regret. This gives a unified geometric framework for understanding the deviation classes controlled by these algorithms and reveals that different first-order methods can control genuinely different classes of deviations. These deviation classes have direct consequences for learning, particularly in games. We study the equilibrium notions induced by exact-form deviations and introduce conservative correlated equilibrium, reflecting both the conservative geometry of the underlying displacement fields and the restricted family of deviations available to the players. We characterize its relation to correlated equilibrium, determine when the resulting equilibrium notions coincide and when they separate, and show how these relationships depend on the geometry and the learning algorithm. Overall, this work gives a unified geometric account of what first-order online learning algorithms are no-regret with respect to, beyond fixed comparators.