arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

图流中聚类系数分布的单遍估计

Single-Pass Estimation of the Clustering Coefficient Distribution in Graph Streams

Cristian Boldrin, C. Seshadhri

arXiv 2609.23489首次发表:更新:

AI 中文总结

针对图流中聚类系数分布估计问题,提出单遍流式算法BOLIDE,结合多种采样策略,在仅存储少量边的情况下,可证明近似分箱按度聚类系数,并在数十亿边和三角形的数据集上准确高效。

AI 中文摘要

三角形计数是网络分析中最基本的问题之一。鉴于现实世界图的规模巨大,长期以来人们设计了多种小空间流式算法来为该问题提供准确的估计。然而,大多数研究结果集中于估计总三角形数或与各个节点关联的三角形数量。在实践中,人们通常需要细粒度的信息来理解三角形的分布,这可以通过聚类系数来刻画。特别是,一项标准的网络分析任务要求计算分箱的按度聚类系数分布,该分布提供了图结构的丰富且信息量大的概要。在这项工作中,我们提出了BOLIDE,这是第一个在流式环境中估计分箱按度聚类系数的有效且实用的算法。我们的算法对边流进行单遍扫描,并且仅允许存储总边数的一小部分。BOLIDE仔细结合了不同的采样策略,以高效地收集跨节点集合的度数和三角形信息。因此,我们的算法可证明地近似分箱按度聚类系数,并对使用的内存量提供保证。我们的实验评估表明,BOLIDE能够准确估计聚类系数分布,同时高效处理具有数十亿条边和三角形的数据集。

英文摘要

Triangle counting is one of the most fundamental problems in network analysis. Given the massive sizes of real-world graphs, there is a long history of small-space streaming algorithms providing accurate estimates for this problem. However, most of the results focus on estimating the total triangle count or the number of triangles incident to individual nodes. In practice, one often wants fine-grained information to understand how triangles are distributed, as captured by clustering coefficients. In particular, a standard network analysis task requires computing the binned degree-wise clustering coefficient distribution, which provides a rich and informative summary of the structure of the graph. In this work we present BOLIDE, the first efficient and practical algorithm for estimating binned degree-wise clustering coefficients in streaming. Our algorithm makes a single pass over the edge stream, and is allowed to store only a small fraction of the total number of edges. BOLIDE carefully combines different sampling strategies to efficiently gather degree and triangle information across sets of nodes. As a result, our algorithm provably approximates the binned degree-wise clustering coefficients, and provides guarantees on the amount of memory used. Our experimental evaluation shows that BOLIDE accurately estimates clustering coefficient distributions while efficiently processing large datasets with billions of edges and triangles.

CommentsVLDB 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑