arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高维中梯度EM学习高斯混合模型是否必须满足$\sqrt{d}$的分离条件?

Is $\sqrt{d}$ Separation Necessary for Gradient EM to Learn Gaussian Mixtures in High Dimensions?

Yiran Zhang, Mo Zhou, Weihang Xu, Maryam Fazel, Simon S. Du

arXiv 2610.07551首次发表:更新:

发表机构

University of California, Berkeley; University of Washington; Amazon, Inc.(加州大学伯克利分校; 华盛顿大学; 亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文证明在高维设置下,梯度EM学习高斯混合模型所需的真实分量间最小分离度必须达到$\Omega(\sqrt{d})$量级,且任何更小的分离度(如$\Omega(d^{0.5-\epsilon})$)在最坏情况下均不足以保证随机初始化下的全局收敛,从而建立了几乎最优的下界。

AI 中文摘要

使用期望最大化(EM)算法及其基于梯度的变体学习高斯混合模型(GMM)是机器学习中的一个基本问题。已知在精确参数化设置下(即分量数量与真实GMM的分量数量相匹配),随机初始化的(梯度)EM无法学习多分量GMM。最近,在过参数化设置下(即使用更多分量),只要真实分量之间具有良好的分离性,梯度EM的全局收敛性已被建立。特别地,真实分量之间的最小分离度需要达到$\Omega(\sqrt{d})$的量级,其中$d$是维度。在本文中,我们证明这种维度依赖性在高维设置中是不可避免的。具体而言,我们考虑一种混合EM算法,该算法对混合权重使用标准EM更新,对分量均值使用梯度EM更新。对于任意$\epsilon > 0$,我们证明当维度足够大时,在最坏情况下,分离度为$\Omega(d^{0.5-\epsilon})$量级不足以保证随机初始化下群体梯度EM在次指数时间内全局收敛,即使在过参数化机制下也是如此。我们的结果建立了通过梯度EM在高维中学习高斯混合所需的真实分离度的几乎最优最坏情况下的下界。

英文摘要

Learning Gaussian mixture models (GMMs) using the Expectation-Maximization (EM) algorithm and its gradient-based variants is a fundamental problem in machine learning. It is known that randomly initialized (gradient) EM fails to learn multi-component GMMs in the exact-parameterized setting, where the number of components matches that of the ground-truth GMM. Recently, global convergence of gradient EM has been established in the over-parameterized setting, where more components are used, provided that the ground-truth components are well separated. In particular, the minimum separation between ground-truth components is required to scale as $Ω(\sqrt{d})$, where $d$ is the dimension. In this paper, we show that this dimensional dependence is unavoidable in high-dimensional settings. Specifically, we consider a hybrid EM algorithm that uses standard EM updates for the mixing weights and gradient EM updates for the component means. For any $ε> 0$, we prove that when the dimension is sufficiently large, in the worst case a separation of order $Ω(d^{0.5-ε})$ is insufficient to guarantee global convergence of population gradient EM in sub-exponential time under random initialization, even in the over-parameterized regime. Our result establishes an almost optimal worst-case lower bound on the ground-truth separation required for learning Gaussian mixtures via gradient EM in high dimensions.

Comments51 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑