发表机构
Laboratoire d'Informatique Gaspard Monge, Université Gustave Eiffel, CNRS(加斯帕尔·蒙日计算机科学实验室,巴黎-塞纳河谷大学,法国国家科学研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明在Wasserstein梯度漂移场下,用Gauss-Newton方案替代梯度下降等价于参数空间中的自然梯度下降,从而在凸性假设下为生成模型训练提供全局收敛保证。
AI 中文摘要
漂移方法最近被引入作为训练生成模型的一种有前景的新范式,迄今已引起广泛关注。在这项工作中,我们聚焦于优化过程如何塑造其收敛性质,揭示可能解释其经验成功的底层机制。更具体地说,当漂移场是Wasserstein梯度时,通过普通欧几里得梯度下降优化漂移损失似乎丧失了利用凸性的任何希望,因此也无法获得任何收敛保证。然而,改变漂移场或优化过程可能导致不同的动态:我们证明,保持Wasserstein梯度作为漂移场,但将梯度下降替换为Gauss-Newton方案,等价于在参数空间中进行自然梯度下降。这一联系使我们能够在适当的凸性假设下为优化过程提供全局收敛保证。
英文摘要
Drifting methods have recently been introduced as a promising new paradigm for training generative models, and have attracted a lot of attention so far. In this work, we focus on how the optimization procedure shapes their convergence properties, shedding light on the underlying machinery that may explain their empirical success. More specifically, when the drift field is a Wasserstein gradient, optimizing the drifting loss via plain Euclidean gradient descent seems to forfeit any hope of exploiting convexity and therefore of obtaining any convergence guarantees. Changing either the drift field or the optimization procedure, however, may lead to different dynamics: we show that keeping the Wasserstein gradient as drift field but replacing gradient descent with a Gauss-Newton scheme amounts to performing a natural gradient descent in parameter space. This connection allows us to provide global convergence guarantees for the optimization procedure under suitable convexity assumptions.