AI 中文总结
该研究针对TDA中聚类问题,提出基于各同调SW核凸组合的核k均值算法,在基准及合成数据集上表现优于基线且计算高效,权重可识别具判别性的同调。
AI 中文摘要
拓扑数据分析(TDA)利用拓扑技术从复杂数据集中提取有意义的基于形状的信息。聚类是数据分析的核心问题,近期人们对理解TDA如何为聚类提供支持产生了浓厚兴趣。现有方法要么在沃尔什斯坦型距离下直接对持久图进行聚类,这在计算上代价高昂;要么使用持久图的向量表示。我们提出一种基于切片沃尔什斯坦(SW)核凸组合的核k均值算法,每个考虑的同调对应一个SW核。与其他持久图的向量表示不同,SW核关于1-沃尔什斯坦距离既稳定又具有判别性。该方法在两个基准数据集上的表现优于基线方法,在第三个合成数据集上仍具有竞争力,同时计算效率高。它还优于所有同调群并集上的单个SW核,以及直接在点云上计算的SW核。凸组合为每个q-同调核分配可解释的权重,我们进一步验证这些权重能够识别具有判别性的同调。
英文摘要
Topological data analysis (TDA) uses topological techniques to extract meaningful shape-based information from complex datasets. Clustering is a central problem in data analysis, and there has been considerable recent interest in understanding how TDA can inform it. Existing approaches either cluster persistence diagrams directly under Wasserstein-type distances, which is computationally expensive, or use vector representations of diagrams. We propose a kernel $k$-means algorithm built on a convex combination of sliced Wasserstein (SW) kernels, one for each homology under consideration. Unlike other vector representations of persistence diagrams, the SW kernel is both stable and discriminative with respect to the $1$-Wasserstein distance. The method outperforms the baselines on two benchmark datasets and remains competitive on a third synthetic dataset, while being computationally efficient. It also outperforms both a single SW kernel on the union of all homology groups and an SW kernel computed directly on the point clouds. The convex combination assigns an interpretable weight to each $q$-homology kernel. We further validate that the weights identify the discriminating homology.
Comments20 Pages, 5 figures