发表机构
Kathmandu University(加德满都大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Sector-Mean,一种基于角扇区分割的O(N)确定性K-Means初始化方法,在保持与K-Means++和Max-Min相当的聚类质量下,显著降低初始化时间并减少迭代次数。
AI 中文摘要
K-Means是最广泛使用的聚类算法之一,但其对初始质心选择的敏感性仍是其收敛速度和聚类准确性的主要瓶颈。本文提出了Sector-Mean初始化,一种具有O(N)时间复杂度的确定性初始化策略,该策略围绕全局质心将二维数据空间划分为角扇区,并使用扇区均值初始化质心。我们在已建立的二维基准(SIPU、Birch)和多个真实世界数据集上评估了该方法,在相同的Lloyd迭代条件下与随机、K-Means++和Max-Min初始化进行了比较。Friedman检验(p<0.05)和Nemenyi事后比较的统计分析表明,在提供与K-Means++和Max-Min相当的聚类质量的同时,Sector-Mean具有显著的计算效率。实验结果表明,与K-Means++和max-min相比,Sector-Mean分别将初始化时间减少了74.9%和59.8%。并且,它产生了最低的平均迭代次数,比K-Means++少约5%,比max-min少16%。这些结果突显了Sector-Mean初始化提供了一种确定性的、计算高效的初始化策略,同时保持了聚类质量。
英文摘要
K-Means is one of the most widely used clustering algorithms, but its susceptibility to initial centroid selection remains a primary bottleneck for its convergence speed and clustering accuracy. This paper proposes Sector-Mean Initialization, a deterministic initialization strategy with O(N) time complexity that partitions the two-dimensional data space into angular sectors around the global centroid and initializes centroids using sector-wise means. We evaluate the method on established two-dimensional benchmarks (SIPU, Birch) and multiple real-world datasets, comparing against random, K-Means++, and Max-Min initialization under identical Lloyd iterations. The statistical analysis of Friedman's test (p<0.05) and Nemenyi post-hoc comparison indicates that, while delivering equivalent clustering quality as K-Means++ and Max-Min, Sector-Mean offers significant computational efficiency. Experimental results show that Sector-Mean reduces the initialization time by 74.9% and 59.8% in comparison to K-Means++ and max-min, respectively. And, it yields the lowest average number of iterations, achieving approximately 5% fewer iterations than K-Means++ and 16% fewer than max-min. These results highlight that Sector-Mean initialization offers a deterministic and computationally efficient initialization strategy while preserving cluster quality.
Comments7 pages