可扩展三角形计数:阈值算法
Scalable Triangle Counting: The Threshold Algorithm
浏览论文内容
中文总结 AI 辅助
提出一种无需先验估计的阈值算法,通过停止规则自适应选择前缀长度,实现随机顺序边流上单遍三角形计数的近似,内存高效且误差可控。
中文摘要 AI 辅助
我们研究随机顺序边流上的单遍三角形计数。我们提出一种极其简单的算法——从流中读取边,直到前缀中观察到 $Q$ 个三角形,然后输出 $Q\\,(m/S)^3$,其中 $S$ 是停止长度——并证明,当与任何边关联的最大三角形数满足 $\eta \le T^{2/3}$ 时,这是 $T$ 的 $(1\pm\varepsilon)$ 近似,概率为 $1-\delta$,使用 $O(\varepsilon^{-2}\log(1/\delta)\\, m/T^{1/3})$ 内存。关键在于,该算法不需要任何 $T$ 的先验估计,这与最先进的基于采样率的算法(McGregor 和 Vorotnikova,PODS 2020;Tsourakakis 等人,KDD 2009)形成鲜明对比。它也不需要预设的内存预算:停止规则自行选择前缀长度,并且可以在读取整个流之前返回估计值。证明基于 Schudy--Sviridenko 集中不等式,用于独立边采样估计器,并结合算法产生的无放回前缀。在六个真实时间流上,该算法的停止前缀遵循预测的立方根缩放,在 $10\\%$ 前缀处实现至多 $6\\%$ 的误差,而无需使用 $T$。在固定的存储边预算下,方差缩减的蓄水池采样器通常更准确,但仅在读取整个流之后。在一个单独的、更大的 $1.8\times10^9$ 边图上,阈值算法读取流的 $0.46\\%$ 并返回 $3.8\\%$ 的误差,而最强的蓄水池基线在墙钟时间上限内未能完成一遍遍历。
英文摘要
We study one-pass triangle counting on random-order edge streams. We present a remarkably simple algorithm---read edges from the stream until $Q$ triangles are observed in the prefix, then output $Q\,(m/S)^3$ where $S$ is the stopping length---and prove that, when the maximum number of triangles incident to any edge satisfies $η\le T^{2/3}$, this is a $(1\pm\varepsilon)$-approximation of $T$ with probability $1-δ$ using $O(\varepsilon^{-2}\log(1/δ)\, m/T^{1/3})$ memory. Crucially, the algorithm does not need any a priori estimate of $T$, in sharp contrast with state-of-the-art sampling-rate based algorithms (McGregor and Vorotnikova, PODS 2020; Tsourakakis et al., KDD 2009). It also does not need a prescribed memory budget: the stopping rule self-selects the prefix length and can return an estimate before reading the entire stream. The proof rests on a Schudy--Sviridenko concentration argument for an independent-edge-sampling estimator, coupled to the without-replacement prefix produced by the algorithm. On six real temporal streams, the algorithm's stopping prefix follows the predicted cube-root scaling and achieves at most $6\%$ error at a $10\%$ prefix, without using $T$. At a fixed stored-edge budget, variance-reduced reservoir samplers are often more accurate, but only after reading the entire stream. On a separate, much larger, $1.8\times10^9$-edge graph, the threshold algorithm reads $0.46\%$ of the stream and returns $3.8\%$ error, while the strongest reservoir baselines do not finish a pass within the wall-clock cap.
发表机构
- Yale University(耶鲁大学)
- University of Massachusetts(马萨诸塞大学)
机构由 AI 辅助整理,请以论文原文为准。