发表机构
University of Macau; Fudan University; Shanghai Jiao Tong University(澳门大学; 复旦大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过误差界性质证明MCR$^2$目标在局部最大化器邻域满足正则性条件,从而保证一阶方法线性收敛,并实验验证其改善表示几何且保持下游准确率。
AI 中文摘要
虽然最大编码率降低(MCR$^2$)目标已成为学习紧凑且判别性表示的有用原则,但其优化几何的理论理解仍然有限,特别是关于为什么一阶方法在其局部最大化器附近往往表现出快速收敛。在这项工作中,我们证明了MCR$^2$目标在其局部最大化器的邻域内满足误差界性质。该性质保证在邻域内任意点,到局部最大化器集合的距离可以被梯度范数上界所界定,直接暗示了诸如Polyak--Łojasiewicz不等式和二次增长等正则性条件。利用这一性质,我们证明在温和条件下,一阶方法线性收敛到MCR$^2$目标的局部最大化器。数值实验验证了预测的局部收敛行为,并提供了经验证据表明,优化MCR$^2$目标改善了表示几何,同时保持了具有竞争力的下游准确率。
英文摘要
While the maximal coding rate reduction (MCR$^2$) objective has become a useful principle for learning compact and discriminative representations, a theoretical understanding of its optimization geometry remains limited, especially regarding why first-order methods often exhibit fast convergence near its local maximizers. In this work, we show that the MCR$^2$ objective satisfies an error-bound property in a neighborhood of its local maximizers. This property guarantees that the distance to the local maximizer set can be bounded above by the gradient norm at any point in the neighborhood, directly implying regularity conditions such as the Polyak--Łojasiewicz inequality and quadratic growth. Leveraging this property, we prove that first-order methods converge linearly to a local maximizer of the MCR$^2$ objective under mild conditions. Numerical experiments validate the predicted local convergence behavior and provide empirical evidence that optimizing the MCR$^2$ objective improves representation geometry while maintaining competitive downstream accuracy.
CommentsThis paper is accepted by NeurIPS 2026