arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

k-means++中中心数量的随机化

Randomizing the Number of Centers in k-means++

Vaclav Rozhon

arXiv 2607.26202首次发表:更新:

AI 中文总结

针对k-means++固定中心数时近似比的问题,该研究将中心数k在K到2K-1间均匀随机选取,证明此预算平滑设置下k-means++可达到常数概率的O(1)近似。

AI 中文摘要

k-means++算法是k-means聚类的标准且广泛使用的初始化方法,但对于固定的中心数k,其最坏情况的期望近似比为Θ(log k)。我们考虑该算法在对手先固定数据集和某个K的情况下,中心数k从{K,…,2K-1}中均匀选取。我们证明,在这种预算平滑设置下,k-means++以常数概率达到O(1)近似。

英文摘要

The $k$-means++ algorithm is a standard and widely used seeding method for $k$-means clustering, but for a fixed number $k$ of centers its worst-case expected approximation ratio is $Θ(\log k)$. We consider the same algorithm when an adversary first fixes the dataset and some $K$; the number of centers $k$ is then chosen uniformly from $\{K,\ldots,2K-1\}$. We prove that $k$-means++ is an $O(1)$-approximation with constant probability in this budget-smoothed setup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑