arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用三元队列寻找多个重叠的稠密子图

Finding Many Overlapping Dense Subgraphs Using Triadic Cohorts

Sabyasachi Basu, C. Seshadhri

arXiv 2609.04890首次发表:更新:

发表机构

Microsoft Research; UC Santa Cruz(微软研究院; 加州大学圣克鲁兹分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出算法 CohortRecovery,基于三元队列框架,可生成覆盖真实网络大量顶点的重叠稠密子图,在多数据集上性能优于多种图聚类算法,能捕获人工整理的真实簇。

AI 中文摘要

图是社会科学、网络安全、计算机基础设施、生物信息学等领域数据的标准表示形式。典型的真实世界图是稀疏的,意味着平均度较小(在数十的量级,而顶点数量超过数百万)。当从某个来源收集图数据时,一项主要任务是进行数据探索,因此任何“稠密”区域都值得关注,因为它指示了特殊结构。一个重要目标是利用稠密子图覆盖图的很大一部分。现有的寻找多个稠密子图的算法不会产生重叠的输出,因此覆盖范围有限;其他产生重叠簇的方法则无法生成稠密子图。本文的主要目标是开发可证明且实用的方法,这些方法能够提供可覆盖真实世界网络很大一部分的重叠稠密子图。我们的贡献是算法 CohortRecovery。我们首先开发了三元队列的数学框架,该框架捕获了可能重叠的“可检测稠密子图”的概念。我们证明,CohortRecovery 可以输出一组稠密子图,使得每个三元队列几乎完全包含在某个稠密子图中。我们给出了 CohortRecovery 的实用实现,并在多种数据集上对其进行了验证。即使在具有数千万条边的图上,它在普通机器上的运行时间通常也不到十分钟。对于众多数据集,CohortRecovery 能够在密度超过 0.8 的非平凡子图中覆盖超过 25% 的顶点,并且显著优于多种可扩展的图聚类/社区检测算法。此外,我们证明 CohortRecovery 的输出捕获了通过人工整理获得的真实簇。

英文摘要

Graphs are a standard representation for data in the social sciences, cybersecurity, computer infrastructure, bioinformatics, and more. Typical real-world graphs are sparse, meaning the average degree is small (in the tens, while the number of vertices is more than millions). When graph data is collected from a source, a major task is to perform data exploration. Thus, any region of ``density" is of interest, since it indicates special structure. An important goal is to cover a significant portion of the graph using dense subgraphs. Existing algorithms that find many dense subgraphs do not have overlapping output, and hence provide limited coverage. Other methods that produce overlapping clusters do not generate dense subgraphs. The main goal of this paper is to develop provable and practical methods that can provide overlapping dense subgraphs that can cover large portions of real-world networks. Our contribution is an algorithm CohortRecovery that achieves this goal. We first develop a mathematical framework of triadic cohorts that captures the notion of ``detectable dense subgraphs'' that potentially overlap. We prove that CohortRecovery can output a set of dense subgraphs, such that each triadic cohort is almost completely contained in some dense subgraph. We give a practical implementation of CohortRecovery and demonstrate it on a variety of datasets. It typically runs in under ten minutes on a commodity machine even on graphs with tens of millions of edges. For numerous datasets, CohortRecovery is able to cover more than 25\% of the vertices in non-trivial subgraphs of density more than 0.8, and is significantly better than a wide variety of scalable graph clustering/community detection algorithms. Moreover, we demonstrate that output of CohortRecovery captures ground truth clusters obtained by manual curation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑