发表机构
University of Waterloo; New York University(滑铁卢大学; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该论文解决了流式集合覆盖问题的复杂度,建立了空间、遍历次数与近似比之间的最优权衡,证明任何近似算法所需空间的下界,并完全确定了该问题的复杂度。
AI 中文摘要
在流式集合覆盖问题中,来自大小为 n 的宇宙的 m 个集合在流中逐一到达,算法被允许使用一次或几次遍历以及 o(mn) 的空间(该空间相对于输入大小是次线性的)来处理流。目标是在最后一次遍历结束时确定覆盖宇宙所需的最少(或近似最少)集合数量。多年来,该问题已被广泛研究,并取得了快速进展,产生了多个在 \tilde{O}(mn^{1/p}) 空间和 O(p) 遍历下的 O(\log{n}) 近似算法。然而,尽管没有任何下界排除在 O(m) 空间和仅两次遍历下实现甚至 O(\log{n}) 近似,过去十年中这方面的进展基本停滞。我们通过为该问题建立最优的三方空间-遍数-近似权衡来为这种缺乏进展提供简单解释:对于流式集合覆盖,任何 \alpha 近似算法在 \alpha \ll n^{1/(p+1)} 时需要 \widetilde{\Omega}\Big(\frac{m}{\alpha} \cdot \big(\frac{n}{\alpha}\big)^{1/p}\Big) 的空间,在 p 次遍历中。鉴于先前的工作,对于任何 \alpha \geq p,该结果在 p 的常数因子和 n,m 的对数因子内是最优的。我们的界在 \alpha 的范围方面也是最优的,并完全解决了流式模型中这一基本问题的复杂度。该结果的证明(令人惊讶地)简单且非技术性,依赖于从通信复杂度中标准指针追逐问题的一个变体进行的随机化归约,并使用随机集合的基本性质。
英文摘要
In the streaming set cover problem, $m$ sets from a universe of size $n$ are arriving one by one in a stream, and the algorithm is allowed to process the stream using one or a few passes and a space of $o(mn)$, which is sublinear in the input size. The goal is to determine the minimal (or approximately minimal) number of sets that cover the universe at the end of the last pass. This problem has been studied extensively over the years with rapid progress that led to several $O(\log{n})$-approximation algorithms in $\tilde{O}(mn^{1/p})$ space and $O(p)$ passes. However, progress on this front has largely stagnated over the past decade, despite the absence of any lower bounds that rule out even an $O(\log{n})$-approximation in $O(m)$ space and just two passes. We provide a simple explanation for this lack of progress by establishing an optimal three-way space-pass-approximation tradeoff for this problem: any $α$-approximation algorithm for streaming set cover requires $$ \widetildeΩ\Big(\frac{m}α \cdot \big(\frac{n}α\big)^{1/p}\Big) $$ space in $p$ passes whenever $α\ll n^{1/(p+1)}$. In light of prior work, this result is optimal up to constant factors in $p$ and logarithmic factors in $n,m$ for any $α\geq p$. Our bound is optimal with respect to the range of $α$ also, and fully settles the complexity of this fundamental problem in the streaming model. The proof of this result is (surprisingly) simple and non-technical and relies on a randomized reduction from a variant of the standard pointer chasing problem in communication complexity, using elementary properties of random sets.
Comments25 pages, 4 figures; Full version appeared in STOC 2026