AI 中文总结
该研究针对数据流的M点查询问题,提出矩阵分解的流式算法,改进分位数问题的内存下界,证明二进CountSketch在分位数问题中接近最优。
AI 中文摘要
我们定义了数据流中的M点查询问题:给定固定矩阵M,目标是在旋转门更新下维护向量x,对每个查询u返回估计值ŷ_u,满足|y_u−ŷ_u|≤ε∥x∥₁,其中y=Mx。我们证明,若M可分解为M=AB,且A、B有空间高效表示,则存在流式算法,内存使用量为O(ε⁻¹∥A∥_{2→∞}∥B∥_{1→1}+(ε⁻¹∥A∥_{∞→∞}∥B∥_{1→1})^{2/3})字。一个重要特例是下三角全1矩阵,对应加性误差±εn的分位数问题,n为数据库规模。我们的框架推广了Cormode和Muthukrishnan(《J. Algorithms》,2005)的旋转门分位数二进方法,简化并改进了Wang等人(SIGMOD,2013)与Luo等人(VLDB,2016)的最优二进CountSketch算法分析。该方法还与Li等人(VLDB J.,2015)差分隐私中的矩阵机制相关:给定数据库x∈ℝ^U和矩阵M,机制输出Mx的私有近似,隐私-误差权衡由M的矩阵分解范数决定。我们还改进了带删除操作的分位数问题的现有下界,证明其内存下界为Ω(ε⁻¹log U)字;同时证明任何分解都满足∥A∥_{2→∞}∥B∥_{1→1}=Ω((log^{1.5}U)/log log U),该新下界表明,对于分位数问题,二进CountSketch在基于分解的方法中几乎最优。
英文摘要
We define the $M$-point query problem in data streams. Given a fixed matrix $M$, the goal is to maintain a vector $x$ under turnstile updates and answer each query $u$ with an estimate $\widehat{y}_u$ satisfying $|y_u-\widehat{y}_u| \leq \varepsilon \|x\|_1$, where $y=Mx$. We show that if $M$ admits a factorization $M=AB$, where $A$ and $B$ have space-efficient representations, then there is a streaming algorithm using $O(\varepsilon^{-1}\|A\|_{2\rightarrow\infty}\|B\|_{1\rightarrow 1}+(\varepsilon^{-1}\|A\|_{\infty\rightarrow\infty}\|B\|_{1\rightarrow 1})^{2/3})$ words of memory. An important special case is the lower-triangular all-ones matrix, which corresponds to the quantiles problem with additive error $\pm \varepsilon n$, where $n$ is the database size. Our framework generalizes the dyadic approach of Cormode and Muthukrishnan (J. Algorithms, 2005) for turnstile quantiles, and simplifies and improves the analysis of the state-of-the-art dyadic CountSketch algorithms of Wang et al. (SIGMOD, 2013) and Luo et al. (VLDB, 2016). Our approach is also related to the matrix mechanism of Li et al. (VLDB J., 2015) in differential privacy: given a database $x\in\mathbb{R}^U$ and a matrix $M$, the mechanism outputs a private approximation to $Mx$, with the privacy-error tradeoff governed by a matrix factorization norm of $M$. We also improve the prior lower bound for quantiles with deletions, showing a memory lower bound of $Ω(\varepsilon^{-1}\log U)$ words. We also show any factorization has $\|A\|_{2\rightarrow\infty}\|B\|_{1\rightarrow 1} = Ω((\log^{1.5} U) / \log\log U)$. This lower bound is new, and shows that for quantiles, the dyadic CountSketch is nearly optimal amongst factorization-based approaches.