任意超图中超边数量估计的次线性算法
Sublinear Algorithms for Estimating the Number of Hyperedges in Arbitrary Hypergraphs
浏览论文内容
中文总结 AI 辅助
针对任意超图中超边数量估计问题,提出对偶访问模型下的随机次线性算法,获(1+ε)近似且查询复杂度为O(ε⁻²√n + √n log n),并证明其下界为Ω(√n)。
中文摘要 AI 辅助
我们研究使用关于顶点数n的次线性查询来估计任意n顶点超图中超边数量的问题。注意,超边数量m可以是n的指数级。对于k-均匀超图,估计m等价于估计平均顶点度数,这一问题在Barhum的硕士论文(魏茨曼科学研究所,2007)中已被研究,该论文采用标准访问模型:采样随机顶点、查询顶点度数、访问关联超边。Barhum的技术无法扩展到任意超图,且简单的下界示例表明,当超边大小无界时,标准访问模型无法产生强次线性算法。为获得非平凡的次线性界,我们考虑访问模型的自然泛化——对偶访问模型,该模型允许采样随机超边(的标签)、查询边大小、访问超边中的顶点。在该模型下,我们提出一种随机算法,以高概率返回m的(1+ε)近似值,仅需O(ε⁻²√n + √n log n)次查询。作为对该算法的补充,我们证明了几乎匹配的下界,表明任何获得m的常数因子近似的算法都需要Ω(√n)次查询。
英文摘要
We study the problem of estimating the number of hyperedges in an arbitrary $n$-vertex hypergraph using sublinear in $n$ queries. Note that the number of hyperedges, $m$, can be exponential in $n$. For $k$-uniform hypergraphs, estimating $m$ is equivalent to estimating the average vertex degree, a problem studied in Barhum's Master's thesis (Weizmann Inst., 2007) under the standard access model of sampling random vertices, querying vertex degrees, and accessing incident hyperedges. Barhum's techniques do not extend to arbitrary hypergraphs, and simple lower-bound examples show that the standard access model cannot yield strongly sublinear algorithms when hyperedges have unbounded size. To obtain non-trivial sublinear bounds, we consider a natural generalization of the access model called the \emph{dual access model}, which allows sampling (labels of) random hyperedges, querying edge sizes, and accessing vertices in a hyperedge. In this model, we give a randomized algorithm that returns a $(1+\varepsilon)$-approximation to $m$ with high probability, making $O(\varepsilon^{-2}\sqrt{n} + \sqrt{n}\log n)$ queries. Complementing our algorithm, we prove a nearly matching lower bound showing that $Ω(\sqrt{n})$ queries are necessary for any algorithm that obtains a constant factor approximation to $m$.