arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08746cs.LGcs.AIcs.DScs.HC

降维与网络科学相遇:UMAP的kNN图上的意义建构

Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph

Duen Horng Chau, Donghao Ren, Fred Hohman, Dominik Moritz

首次发表
浏览论文内容

中文总结 AI 辅助

研究探索UMAP内部kNN图的潜力,应用PageRank、k核分解和聚类系数等标准图算法增强数据意义建构,经对MNIST和Fashion MNIST评估,证明这些基于图的分析实用且与专门方法有竞争力或互补。

中文摘要 AI 辅助

虽然UMAP广泛用于探索高维数据,但典型工作流程侧重于其低维嵌入,很大程度上忽略了UMAP内部构建的丰富k近邻(kNN)图。该图在UMAP的二维投影引入失真之前,在其原始高维空间中对数据流形进行编码。我们展示了这种内部表示的未开发潜力,表明应用于该图的标准图算法如何增强数据意义建构:(1)PageRank识别代表性数据点,(2)k核分解揭示密集核心区域与稀疏外围,(3)聚类系数检测具有高度相似数据点的紧密邻域。通过对MNIST和Fashion MNIST的定量和定性评估,我们表明这些基于图的分析不仅实用,而且与专门方法(如用于样本选择的k-medoids、用于基于密度聚类的HDBSCAN)具有竞争力或互补性。

英文摘要

While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. This graph encodes the data manifold in its original high-dimensional space, before the distortion that UMAP's 2D projection introduces. We demonstrate the untapped potential of this internal representation, showing how standard graph algorithms applied to this graph enhance data sensemaking: (1) PageRank identifies representative data points, (2) k-core decomposition reveals dense core regions versus sparse periphery, and (3) clustering coefficient detects tight-knit neighborhoods with highly-similar data points. Through quantitative and qualitative evaluation on MNIST and Fashion MNIST, we show that these graph-based analyses are not only practical but also competitive with or complementary to purpose-built methods (e.g., k-medoids for exemplar selection, HDBSCAN for density-based clustering).

发表机构

  • Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑