arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

往返KNN聚类:有向最近邻图上的多尺度层次簇检测

Round-Trip KNN Clustering: multiscale hierarchical cluster detection on directed nearest-neighbour graphs

Eraldo Pereira Marinho, Caetano Mazzoni Ranieri, Fabricio Aparecido Breve

arXiv 2610.06795首次发表:更新:

发表机构

São Paulo State University (UNESP)(圣保罗州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出往返KNN聚类方法,在有向最近邻图上多尺度检测层次簇,无需预设簇数,经细化提升聚类精度,在合成数据上表现优异。

AI 中文摘要

我们提出了往返KNN聚类(RTKNNC),一种基于图的方法,用于在多个邻域尺度上发现簇结构,而无需预先指定簇的数量。与先使k近邻(KNN)图无向化的方法不同,RTKNNC保留邻居关系的两个方向:给定点选择哪些点,以及哪些点选择该点。入向选择被视为加权投票,帮助决定在递归的前向-反向遍历过程中哪些局部连接保持可见。随着K的增加重复该过程,揭示出随着邻域尺度增长,群组如何持续存在或合并;对于结构细化之前的参考逆平方模型,簇可以合并但不能分裂。由于图连通性偶尔会通过稀疏桥或小的重叠区域连接不同的群组,我们添加了一个可选的无需标签的细化步骤。它首先测试已形成的组件是否更好地由两个或三个高斯子总体描述,并且仅当提议的群组足够大且与可见的KNN图一致时才接受细分。在八个合成数据集和K=2,…,16的情况下,独立的C和Python实现在所有120次参考运行中产生了相同的分区。细化将可变密度基准上的调整兰德指数从0.7817提高到0.9627,在稀疏桥基准上从0.8083提高到0.9853。与七种外部聚类方法的比较显示,在保持无标签簇构建过程的同时,性能具有竞争力。

英文摘要

We introduce Round-Trip KNN Clustering (RTKNNC), a graph-based method for finding cluster structure at several neighbourhood scales without requiring the number of clusters in advance. Unlike approaches that first make a $k$-nearest-neighbour (KNN) graph undirected, RTKNNC keeps both directions of the neighbour relation: which points a given point selects and which points select it. Incoming selections are treated as weighted votes that help decide which local connections remain visible during a recursive forward-and-reverse traversal. Repeating the procedure for increasing $K$ reveals how groups persist or merge as the neighbourhood scale grows; for the reference inverse-square model before structural refinement, clusters can merge but do not split. Because graph connectivity can occasionally join distinct groups through a sparse bridge or a small region of overlap, we add an optional label-free refinement. It first tests whether an already formed component is better described by two or three Gaussian subpopulations, and accepts a subdivision only when the proposed groups are large enough and consistent with the visible KNN graph. Across eight synthetic datasets and $K=2,\ldots,16$, independent C and Python implementations produced identical partitions in all 120 reference runs. Refinement increased adjusted Rand index from $0.7817$ to $0.9627$ on a variable-density benchmark and from $0.8083$ to $0.9853$ on a sparse-bridge benchmark. Comparisons with seven external clustering methods show competitive performance while preserving a label-free cluster-construction process.

CommentsSubmitted to Knowledge and Information Systems (KAIS). 32 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑