arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

稀疏主成分分析与鲁棒稀疏估计的快速算法

Fast Algorithms for Sparse PCA and Robust Sparse Estimation

Giannis Iakovidis, Ankit Pensia

arXiv 2609.09701首次发表:更新:

发表机构

University of Wisconsin-Madison; Carnegie Mellon University(威斯康星大学麦迪逊分校; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对稀疏PCA认证问题,提出运行时间为$O(d^2+d k^{O(\log k)})$的双准则算法,并在样本访问模型中实现次二次时间,首次为鲁棒稀疏估计提供二次及次二次算法。

AI 中文摘要

我们研究稀疏PCA认证的快速算法。给定一个半正定矩阵$M$,该问题要求要么排除一个大的$k$-稀疏二次型,要么返回一个高值(松弛)见证。标准的半定松弛提供了此类证书,但现有的通用求解器需要$\Omega(d^4)$时间。我们给出一个双准则算法,运行时间为$O(d^2+d k^{O(\log k)})$:如果某个$k$-稀疏单位向量的二次型大于$2$,它返回一个$O(k^2)$-稀疏单位向量或一个值至少为$1$的SDP可行矩阵。对于$k\leq\exp(O(\sqrt{\log d}))$,该运行时间为$O(d^2)$。我们还在样本访问模型中突破了二次屏障:给定$n=d^{o(1)}$个样本,我们的算法在$d^{2 - \Omega(1)}$时间内获得相关的一侧证书,适用于$k=\mathrm{polylog}(d)$,而无需形成经验协方差矩阵。作为应用,这些证书例程为广泛分布族提供了首个二次和次二次时间的鲁棒稀疏估计算法。我们的稀疏PCA算法将高值稀疏方向约简为大相关图中的有界半径集合,并搜索由此产生的候选支持集。次二次实现利用快速相关检测构建该图。

英文摘要

We study fast algorithms for sparse-PCA certification. Given a positive semidefinite matrix $M$, the problem asks either to rule out a large $k$-sparse quadratic form or to return a high-value (relaxed) witness. The standard semidefinite relaxation provides such certificates, but existing general-purpose solvers require $Ω(d^4)$ time. We give a bicriteria algorithm running in $O(d^2+d k^{O(\log k)})$ time: if some $k$-sparse unit vector has quadratic form greater than $2$, it returns either an $O(k^2)$-sparse unit vector or an SDP-feasible matrix of value at least $1$. For $k\leq\exp(O(\sqrt{\log d}))$, this running time is $O(d^2)$. We also go below the quadratic barrier in the sample-access model: Given $n=d^{o(1)}$ samples, our algorithm obtains a related one-sided certificate in $d^{2 - Ω(1)}$ time for $k=\mathrm{polylog}(d)$, without forming the empirical covariance matrix. As an application, these certificate routines yield the first quadratic and subquadratic-time algorithms for robust sparse estimation for broad families of distributions. Our sparse-PCA algorithm reduces a high-value sparse direction to a bounded-radius set in the graph of large correlations and searches the resulting candidate supports. The subquadratic implementation constructs this graph using fast correlation detection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑