基于密度幂散度测度的鲁棒K-means聚类
Robust K-means Clustering using the Density Power Divergence Measure
- Case Western Reserve University(凯斯西储大学)
- John Carroll University(约翰卡罗尔大学)
- The University of Texas at El Paso(德克萨斯大学埃尔帕索分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出MK-means DPD和DC-MK-means DPD两种鲁棒聚类方法,结合DPD测度与马氏距离,还提出两个鲁棒评估指标,在模拟数据和真实数据集上验证了其优于现有方法的性能。
AI中文摘要:
我们提出了一种鲁棒聚类方法MK-means DPD,该方法结合密度幂散度(DPD)测度与马氏距离来估计聚类中心和协方差矩阵,使其对异常值具有鲁棒性,且能适配异质的椭圆形聚类,这与经典K-means算法不同。由于基于马氏距离的K-means缺乏通用收敛保证,我们进一步提出了收敛变体密度一致MK-means DPD(DC-MK-means DPD),该方法通过逐点DPD损失重新定义聚类分配步骤。我们证明了形式定理,确定所得算法在有限步数内收敛。我们还提出了两个新的鲁棒内部评估指标:中位数戴维斯-布尔丁指数和修剪后的卡林斯基-哈拉巴斯指数,以确保性能比较本身不会被异常值扭曲。所提方法在模拟数据上验证了有效性,显示出优于现有方法的性能;在两个真实数据集上也得到验证:Iris数据用于识别相似物种,以及全球各国的COVID-19病死率和感染率数据,用于研究所得聚类的地理和社会经济模式。
英文摘要:
We introduce a robust clustering method, MK-means DPD, that estimates cluster centers and covariance matrices using density power divergence (DPD) measures combined with Mahalanobis distance, making it resistant to outliers and adaptable to heterogeneous, elliptical clusters, unlike the classical K-means algorithm. Since Mahalanobis distance-based K-means lacks a general convergence guarantee, we further introduce a convergent variant, Density-Consistent MK-means DPD (DC-MK-means DPD), which redefines the cluster assignment step in terms of a pointwise DPD loss. We prove a formal theorem establishing that the resulting algorithm converges in a finite number of steps. We also propose two new robust internal evaluation indices, a Median Davies-Bouldin Index and a Trimmed Calinski-Harabasz Index, to ensure that performance comparisons are not themselves distorted by outliers. The efficacy of the proposed methods is demonstrated on simulated data, showing superiority over existing methods, and on two real datasets: Iris data, to identify similar species, and COVID-19 case fatality rate and infection rate data for countries worldwide, examining the resulting clusters' geographic and socio-economic patterns.