用于Bures-Wasserstein重心的投影黎曼梯度下降:单位步长下与维度无关的线性收敛
Projected Riemannian Gradient Descent for the Bures-Wasserstein Barycenter: Dimension-Independent Linear Convergence at Unit Step Size
浏览论文内容
中文总结 AI 辅助
该研究针对Bures-Wasserstein重心计算的二分性问题,提出投影黎曼梯度下降算法,在单位步长下实现与维度无关的线性收敛,多项式改进了迭代复杂度,还将该保证扩展到了不变矩阵投影场景。
中文摘要 AI 辅助
正定矩阵集合的Bures-Wasserstein(BW)重心计算广泛应用于机器学习、最优传输和量子信息领域。实践中使用的固定点迭代——单位步长下的黎曼梯度下降(RGD)收敛迅速,但现有分析存在二分性:单位步长的保证具有最坏情况下对维度的指数依赖,而与维度无关的保证则需要小步长,这会丧失经验速度。我们解决了该二分性问题,并非通过改进单位步长RGD的保证,而是提出了一种投影RGD算法,该算法在单位步长下实现了与维度无关的线性收敛。所达到的速率为(1 - κ^(-3/2)),其中κ是集合的条件数,这也在迭代复杂度上多项式改进了最佳小步长保证(κ^(3/2)对κ^(5/2))。关键在于一个新颖的投影引理:将正定矩阵的特征值裁剪到区间[α, β]是集合{S: αI ≤ S ≤ βI}上的闭式非扩张(1-Lipschitz)BW度量投影——该表述与其已知的单侧对应物不同,并非由凸性推导而来。此外,该投影是免费的:它复用了下一次迭代无论如何都必须执行的特征分解,因此投影迭代和未投影迭代每一步的成本相同。相同的分析还涵盖了Brahmachari等人(2025)提出的不变矩阵投影问题,我们将其定点算法识别为全测地子流形上的单位步长RGD,从而将与维度无关的保证逐字扩展到该场景。
英文摘要
The computation of the Bures-Wasserstein (BW) barycenter of an ensemble of positive definite matrices arises throughout machine learning, optimal transport, and quantum information. Riemannian gradient descent (RGD) at unit step size -- the fixed-point iteration used in practice -- converges rapidly, yet existing analyses present a dichotomy: unit-step guarantees carry worst-case exponential dependence on the dimension, while dimension-independent guarantees require small step sizes that forfeit the empirical speed. We resolve this dichotomy, not by improving the guarantees for unit-step RGD, but by proposing a Projected RGD algorithm that achieves dimension-independent linear convergence at unit step size. The achieved rate, $(1 - κ^{-3/2})$, where $κ$ is the condition number of the ensemble, also polynomially improves on the best small-step guarantee ($κ^{3/2}$ versus $κ^{5/2}$ iteration complexity). The crux is a novel Projection Lemma: clipping the eigenvalues of a positive matrix to an interval $[α, β]$ is the closed-form, non-expansive (1-Lipschitz) BW-metric projection onto the set $\{S : αI \leq S \leq βI\}$ -- a statement which, unlike its known one-sided counterpart, does not follow from convexity. The projection is moreover free: it reuses an eigendecomposition the next iteration must perform in any case, so the projected and unprojected iterations cost the same per step. The same analysis covers the invariant matrix projection problem of Brahmachari et al. (2025), whose fixed-point algorithm we identify as unit-step RGD on a totally geodesic submanifold, thereby extending the dimension-independent guarantee to that setting verbatim.
发表机构
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。