劳埃德K均值聚类算法是伪装的弗兰克-沃尔夫算法
Lloyd's $K$-Means Clustering Algorithm Is Frank-Wolfe in Disguise
另 1 家 · 查看机构详情
- Old Dominion University(奥多明尼昂大学)
- Miami University(迈阿密大学)
- Université de Montréal(蒙特利尔大学)
- Mila(米拉研究所)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究劳埃德K均值算法与弗兰克-沃尔夫算法的联系,证明前者是后者特殊情况,利用FW方法进展得出SSE目标局部最小值收敛速度,开发处理空聚类的FW变体,通过模拟研究说明结果。
中文摘要 AI 辅助
劳埃德K均值算法,也称为朴素K均值算法,是一种广泛使用的临时优化启发式算法,旨在通过迭代聚类细化最小化数据集所有K划分上的平方误差之和(SSE)。本文建立了劳埃德算法与弗兰克-沃尔夫(FW)算法之间的新联系,证明劳埃德算法是FW的特殊情况。利用FW方法在凹目标上的最新进展,得出SSE目标局部最小值的非渐近O(1/t)收敛速度。为处理空聚类,开发了半光滑目标的FW变体,保留由初始SSE值单独控制的相同收敛速度。通过对球形高斯混合和真实世界图像分割数据集的模拟研究说明了研究结果。
英文摘要
Lloyd's $K$-means algorithm, also known as naïve $K$-means, is a widely used ad hoc optimization heuristic, designed to minimize the sum of squared errors (SSE) across all $K$-partitions of a dataset via iterative cluster refinement. In this work, we establish a novel connection between Lloyd's algorithm and the Frank-Wolfe (FW) algorithm, a prominent first-order method for projection-free optimization. We demonstrate that Lloyd's algorithm is a special case of FW. Leveraging recent advances in FW methods for concave objectives, we derive a non-asymptotic $\mathcal{O}(1/t)$ convergence rate to a local minimum of the SSE objective. To account for empty clusters, an outcome possible under Lloyd's greedy assignment, we develop an FW variant for semismooth objectives while retaining the same convergence rate that is solely controlled by the initial SSE value. We illustrate our findings with a simulation study for spherical Gaussian mixtures and a real-world image segmentation dataset.