arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08243cs.LGstat.ML

一种可解释的k均值++算法的古德-图灵重启准则

An interpretable Good--Turing restart criterion for k-means++

Renato Cordeiro de Amorim

首次发表
浏览论文内容

中文总结 AI 辅助

研究k均值++算法重启次数问题,提出结合古德-图灵估计等的GTRC准则,该准则能依数据集难度调整重启次数,使聚类质量具竞争力,为固定重启次数提供可解释且有原则的替代方案。

中文摘要 AI 辅助

k均值++算法通常会多次重启以避免陷入较差的局部最优解,但重启次数几乎总是任意选择且均匀应用,而不考虑数据集的难度。这破坏了基于这种选择的任何比较,在简单数据集上浪费计算资源,同时可能无法满足困难数据集的需求。我们引入了GTRC,一种结合了古德-图灵估计、已证明的无条件界限以及进一步重启会改善当前结果的概率的基于置信度的界限的重启准则,一旦该概率低于用户指定的容差ε就停止。在36个数据集上,GTRC达到了与精心选择的固定重启次数具有竞争力的聚类质量,而使用的重启次数随数据集难度有很大且适当的变化,由一个可解释的、依赖数据的信号而非固定规则控制。GTRC为预先固定k均值++重启次数提供了一个有原则且可报告的替代方案。

英文摘要

The k-means++ algorithm is commonly restarted multiple times to avoid poor local optima, yet the number of restarts is almost always chosen arbitrarily and applied uniformly regardless of data set difficulty. This undermines any comparison relying on such a choice and wastes computation on easy data sets while potentially under-serving hard ones. Here, we introduce the Good-Turing Restart Criterion (GTRC). This combines a Good-Turing estimate, a proven unconditional bound, and a confidence-based bound on the probability that a further restart would improve on the current result, stopping once this probability falls below a user-specified tolerance. Our experiments on 34 real-world data sets show that GTRC identifies the point beyond which further k-means++ restarts yield only negligible improvement, achieving a more favourable balance between the number of restarts used and clustering quality than three existing stopping rules for multistart local search and popular fixed restart counts. Software: https://github.com/RCdeAmorim/Good-Turing-Restart-Criterion.

发表机构

  • School of Computer Science and Electronic Engineering, University of Essex(计算机科学与电子工程学院,埃塞克斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑