arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31848cs.LGmath.OC

平均镜像下降与对偶梯度方法:熵正则化Gromov-Wasserstein问题的收敛算法

Averaged Mirror Descent and Dual Gradient Methods: Convergent Algorithms for Entropic Gromov-Wasserstein Problems

Joanna Marks, Gabriel Rioux, Riccardo Passeggeri

首次发表
浏览论文内容

中文总结 AI 辅助

针对熵正则化Gromov-Wasserstein问题,提出平均镜像下降算法并证明其对任意成本收敛,同时证明固定步长对偶梯度法的收敛性,在经典MD失败的案例上两者均有效。

中文摘要 AI 辅助

Gromov-Wasserstein(GW)距离度量度量测度(mm)空间之间的差异,并仅基于其内在结构识别它们之间的最优对齐。由于它识别同构的mm空间,它为可能具有同构表示的非同质数据集提供了一种自然的距离概念。为了加速GW距离的计算,许多从业者采用熵正则化来获得熵GW(EGW)问题。最流行的EGW求解器是镜像下降(MD)算法,它将EGW计算简化为迭代过程,其中每次迭代解决一个熵最优传输(EOT)问题。尽管其广泛使用,MD对此问题的收敛性仅针对受限的成本类别建立。另一方面,最近提出的对偶梯度方法适用于一般成本,但需要选择依赖于正则化参数的步长。为解决这两个问题,我们引入了平均镜像下降(AMD),它平均连续的MD步骤,并证明了其对任意成本的收敛性。然后,我们建立了具有固定步长的对偶梯度方法也以更复杂的迭代为代价对任意成本收敛。在两种情况下,我们还考虑了在实践中不可避免的不精确迭代。我们比较了这些方法在各种设置下的经验性能,特别是表明AMD和对偶梯度方法在经典MD失败的示例上均收敛。

英文摘要

The Gromov-Wasserstein (GW) distance measures the discrepancy between metric measure (mm) spaces and identifies optimal alignments between them based solely on their intrinsic structure. Since it identifies isomorphic mm spaces, it provides a natural notion of distance for heterogeneous datasets which may admit isomorphic representations. In order to accelerate computation of GW distances, many practitioners employ entropic regularization to obtain an Entropic GW (EGW) problem. The most popular EGW solver is the Mirror Descent (MD) algorithm, which reduces EGW computations to an iterative process where an entropic optimal transport (EOT) problem is solved at each iteration. Despite its widespread use, the convergence of MD for this problem has only been established for restricted classes of costs. On the other hand, a recently proposed dual gradient method is available for general costs, but requires a choice of step size which depends on the regularization parameter. To address these two issues, we introduce Averaged Mirror Descent (AMD), which averages consecutive MD steps, and prove its convergence for arbitrary costs. Then, we establish that the dual gradient method with a fixed step size also converges for arbitrary costs at the cost of a more complicated iteration. In both cases, we also account for inexact iterations which are inescapable in practice. We compare the empirical performance of these methods across various settings and, in particular, show that AMD and the dual gradient method both converge on an example where classical MD fails.

↑