发表机构
CNRS; Sorbonne Université(法国国家科学研究中心; 索邦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SK-Wasserstein距离用于持久态图,其计算复杂度为O(NlogN),在实验中相比W2近似方法实现超2000倍加速,且在聚类任务中表现优于W2的平均链接划分,可兼容多种机器学习方法。
AI 中文摘要
本文提出了Sierpiński-Knopp(SK)Wasserstein距离,这是一种适用于持久态图的快速度量。SK-Wasserstein距离记为\boldsymbol{d}_{\text{SK}},它通过上对角三角形上的Sierpiński-Knopp空间填充曲线,将图中的点及其在对角线上的投影映射到单位区间。随后,编码后的点集通过一维最优匹配以\boldsymbol{O}(N\boldsymbol{\text{log}}\boldsymbol{N})的步骤进行高效匹配,从而在两个输入持久态图之间生成显式的感知对角点分配。研究表明,SK-Wasserstein距离可控制图之间的经典2-Wasserstein距离,允许显式的等距嵌入到希尔伯特空间中,并诱导出正定高斯核,使得所得几何结构可直接与基于欧氏距离和核的学习方法兼容。基于沿该曲线的点分配,还提出了更紧密的替代不相似度\boldsymbol{W}_\boldsymbol{\boldsymbol{\text{Γ}}}。在包含227个图的12个科学数据集上进行的实验显示,与最先进的\boldsymbol{W}_2近似方法相比,\boldsymbol{d}_{\text{SK}}在每个数据集上的中位数加速比为626倍,在整个基准测试中的总加速比为2100倍。由\boldsymbol{d}_{\text{SK}}和\boldsymbol{W}_\boldsymbol{\boldsymbol{\text{Γ}}}得到的平均链接划分,分别在12个数据集中的8个上与对应的\boldsymbol{W}_2划分完全匹配。基于\boldsymbol{d}_{\text{SK}}的希尔伯特k-均值和高斯谱聚类,相对于基准参考划分,分别获得了0.756和0.800的平均调整兰德指数(ARI),而\boldsymbol{W}_2的平均链接划分获得的ARI为0.750。实验中还展示了高斯\boldsymbol{d}_{\text{SK}}核可用于有序图集合的连续分割等其他基于核的分析任务。
英文摘要
This paper introduces the Sierpiński-Knopp (SK) Wasserstein distance, a fast metric between persistence diagrams. The SK-Wasserstein distance, denoted $d_{\mathrm{SK}}$, maps diagram points and their diagonal projections to the unit interval via the Sierpiński-Knopp space-filling curve on the upper diagonal triangle. The encoded point sets are then efficiently matched via one-dimensional optimal assignment, in \(O(N\log N)\) steps, yielding an explicit diagonal-aware point assignment between the two input persistence diagrams. We show that the SK-Wasserstein distance controls the classical \(2\)-Wasserstein distance between diagrams, admits an explicit isometric embedding into a Hilbert space, and induces a positive-definite Gaussian kernel, making the resulting geometry directly compatible with Euclidean and kernel-based learning methods. A tighter surrogate dissimilarity, noted \(W_Γ\), is also introduced based on the point assignments along the curve. Experiments on 12 scientific collections comprising 227 diagrams show median per-collection speedup of \(d_{\mathrm{SK}}\) over state-of-the-art approximations of \(W_2\) is \(626\times\), while the aggregate speedup over the full benchmark is \(2100\times\). Average-linkage partitions obtained from \(d_{\mathrm{SK}}\) and \(W_Γ\) each exactly match the corresponding \(W_2\) partition on 8 of the 12 collections. Hilbert \(k\)-means and Gaussian spectral clustering, both based on \(d_{\mathrm{SK}}\), achieve mean adjusted Rand indices (ARI) of \(0.756\) and \(0.800\), respectively, with respect to the benchmark reference partitions, compared to \(0.750\) obtained by average linkage on \(W_2\). The Gaussian \(d_{\mathrm{SK}}\) kernel supports other kernel-based analysis tasks, as illustrated by its use for contiguous segmentation of ordered diagram collections in our experiments.
Comments49 pages, 11 figures, 6 tables. Code and reproducibility package: https://github.com/sebastien-tchitchek/SK-Wasserstein-Reproducibility