AI 中文总结
该研究针对$\u2113_p$子空间逼近强核心集问题,分别在1≤p<2和p>2两种情形下通过不同技术改进了核心集大小的ε依赖界至ε⁻²量级,且算法运行效率较高,部分结果接近采样下界。
AI 中文摘要
我们研究$\u2113_p$子空间逼近的强核心集。给定矩阵$A\in\mathbb{R}^{n\times d}$,目标是采样并重新缩放其少量行得到$SA$,使得对每个维数不超过$k$的子空间$F\subseteq\mathbb{R}^d$,同时满足$\left\\|SA(I-P_F)\right\\|_{p,2}^p=(1\pm\varepsilon)\left\\|A(I-P_F)\right\\|_{p,2}^p$,其中$P_F$是到$F$的正交投影算子。Woodruff和Yasuda [WY25](FOCS 2025)得到的核心集大小在$1\leq p<2$时为$\widetilde{O}_p(k\varepsilon^{-4/p})$,在$p>2$时为$\widetilde{O}_p(k^{p/2}\varepsilon^{-p})$。我们将这些界分别改进为$\widetilde{O}_p(k\varepsilon^{-2})$和$\widetilde{O}_p(k^{p/2}\varepsilon^{-2})$。对于$1\leq p<2$,我们的算法运行时间为$\widetilde{O}_p(\mathrm{nnz}(A)+d^\omega+k\varepsilon^{-2})$。当$k+1\geq C\log(1/\varepsilon)$($C$为绝对常数)时,所得核心集大小在对数因子意义下匹配采样下界[LWW21]。对于$p>2$,我们的算法运行时间为$\widetilde{O}_p(\mathrm{nnz}(A)+d^\omega)$,与Woodruff-Yasuda框架的运行时间一致。\n我们在两种情形下使用了不同技术。对于$1\leq p<2$,我们将双准则低秩拆分与Lewis权重采样以及与输出维数无关的经验过程界相结合。对于$p>2$,我们对Woodruff-Yasuda构造给出了更精细的分析。通过在整个行数递推过程中保留其采样概率中的截断项,我们证明它实现了改进的$\varepsilon^{-2}$依赖关系。
英文摘要
We study strong coresets for $\ell_p$ subspace approximation. Given a matrix $A\in\mathbb{R}^{n\times d}$, the goal is to sample and rescale a small number of its rows to obtain $SA$ such that $\left\|SA(I-P_F)\right\|_{p,2}^p=(1\pm\varepsilon)\left\|A(I-P_F)\right\|_{p,2}^p$ simultaneously for every subspace $F\subseteq\mathbb{R}^d$ of dimension at most $k$, where $P_F$ is the orthogonal projector onto $F$. Woodruff and Yasuda [WY25] (FOCS 2025) obtained coreset sizes $\widetilde{O}_p(k\varepsilon^{-4/p})$ for $1\leq p<2$ and $\widetilde{O}_p(k^{p/2}\varepsilon^{-p})$ for $p>2$. We improve these bounds to $\widetilde{O}_p(k\varepsilon^{-2})$ and $\widetilde{O}_p(k^{p/2}\varepsilon^{-2})$, respectively. For $1\leq p<2$, our algorithm runs in $\widetilde{O}_p(\mathrm{nnz}(A)+d^ω+k\varepsilon^{-2})$ time. The resulting coreset size matches the sampling lower bound [LWW21] up to logarithmic factors when $k+1\geq C\log(1/\varepsilon)$ for an absolute constant $C$. For $p>2$, our algorithm runs in $\widetilde{O}_p(\mathrm{nnz}(A)+d^ω)$ time, matching the running time of the framework of Woodruff and Yasuda. We use different techniques in the two regimes. For $1\leq p<2$, we combine a bicriteria low-rank split with Lewis-weight sampling and empirical-process bounds independent of the output dimension. For $p>2$, we give a sharper analysis of the Woodruff-Yasuda construction. By retaining the truncation in its sampling probabilities throughout the row-count recurrence, we show that it achieves the improved $\varepsilon^{-2}$ dependence.