发表机构
Northeastern University; MIT CSAIL; Google Research(东北大学; 麻省理工学院计算机科学与人工智能实验室; 谷歌研究)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对偏序集的单调性测试问题,提出一种在√n^(1+o(1))次遍历中实现(2+ε)近似的流式算法,首次关联亚线性时间与流式最大匹配估计,改进了现有算法的遍历复杂度。
AI 中文摘要
考虑一个偏序集——等价于一个含n个顶点的有向无环图(DAG)G=(V, E)——以及其顶点集上的布尔函数f: V→{0,1}。若对所有(u, v)∈E都有f(u)≤f(v),则称f是单调的。尽管关于单调性测试的查询复杂度已有大量研究文献,我们转而关注其空间复杂度,并在流式场景中开启该问题的研究。具体而言,G的边以任意顺序到达,目标是使用Õ(n)空间估计给定函数f与单调性的距离。需注意,该空间虽允许接收和存储f,但远小于可能包含多达Ω(n²)条边的输入图G。我们的主要结果是一种算法,可在√n^(1+o(1))次遍历中对单调性距离进行(2+ε)近似。我们还证明,对于任何O(1)近似,这是所能期望的最优遍历复杂度,除非改进st-可达性的最新流式算法——这是一个被广泛研究的问题。在技术层面,我们的算法近似计算G的传递闭包(的子图)中的最大匹配规模。尽管最大匹配问题在流式场景中已受到大量关注,但在传递闭包中计算它需要截然不同的思路。事实上,我们工作的一项主要贡献是首次将用于估计最大匹配规模的亚线性时间算法与流式场景关联起来。现有现成的亚线性时间算法在我们的场景中仅能得到n√n^(1+o(1))次遍历的算法,我们通过允许更强的查询(如顶点查询和子集查询)大幅改进了这些算法,这些查询对于我们的问题而言,其实现效率与更标准的邻接矩阵查询和列表查询相当。
英文摘要
Consider a poset - or equivalently an $n$-vertex DAG $G=(V, E)$ - and a boolean function $f: V \rightarrow \{0, 1\}$ on its vertex set. We say $f$ is monotone if $f(u) \leq f(v)$ for all $(u, v) \in E$. While there is extensive literature on the query complexity of testing monotonicity, we focus instead on the space complexity and initiate the study of this problem in the streaming setting. Namely, the edges of $G$ arrive in an arbitrary order, and the goal is to estimate distance to monotonicity of a given function $f$ using $\widetilde{O}(n)$ space. Note that while this space allows receiving and storing $f$, it is much smaller than the input graph $G$ which could have up to $Ω(n^2)$ edges. Our main result is an algorithm that $(1+ε)$-approximates distance to monotonicity in $\sqrt{n}^{1+o(1)}$ passes. We also prove that this is the best pass-complexity one can hope for, for any $O(1)$-approximation, short of improving the state-of-the-art streaming algorithm for $st$-reachability, which is a very well-studied problem. On the technical side, our algorithm approximates the size of maximum matching in (a subgraph of) the transitive closure of $G$. While the maximum matching problem has received significant attention in the streaming setting, the fact that we are computing it in the transitive closure requires very different ideas. In fact, a main contribution of our work is to connect sublinear time algorithms for estimating the maximum matching size to the streaming setting for the first time. While existing off-the-shelf sublinear time algorithms only result in an $n\sqrt{n}^{1+o(1)}$ pass algorithm in our setting, we show how to significantly improve upon them by allowing stronger queries (such as vertex and subset queries) that can be implemented just as efficiently as more standard adjacency matrix and list queries for our problem.
Comments42 pages, 1 figure