arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23332stat.MEmath.STstat.TH

流式主成分分析:从几何视角进行平均化

Streaming PCA: averaging from a geometric perspective

Tuan Pham, Alessandro Rinaldo, Purnamrita Sarkar

首次发表
浏览论文内容

中文总结 AI 辅助

本文从几何视角研究流式PCA,通过外空间等价性建立Oja算法的收敛保证,提出自适应平均化估计量,无需预知特征间隙即可达到近最优收敛率,并应用于椭圆成分分析。

中文摘要 AI 辅助

我们研究了内存约束下的主成分分析(PCA),这一设置在大规模数据分析中日益重要。我们的重点是Oja算法,它是一种单遍、内存高效的算法,在秩一情形下仅需$O(p)$存储,在$k$-PCA情形下需$O(pk)$存储。主要目标是开发一种基于Oja算法的程序,该程序在未知特征间隙的情况下达到最优,而特征间隙通常是有效调整学习率所必需的。为此,我们首先引入关于$k$-PCA收敛的几何视角:通过将Oja迭代嵌入外空间,我们建立了环境空间中的$k$-PCA与外空间中的$1$-PCA之间的精确等价性。这一几何视角使我们能够建立若干收敛保证,包括仅基于单一学习率调度的保证。基于这一几何视角,我们为$k$-PCA开发了平均化理论,并证明所得的平均化估计量是自适应的:它在未知特征间隙的情况下达到近乎最优的收敛率。作为应用,我们为椭圆成分分析(ECA)构建了内存高效的估计量。我们进行了模拟研究和真实数据分析,以展示所提出算法的优势。

英文摘要

We study principal component analysis (PCA) under memory constraints, a setting that is increasingly important in large-scale data analysis. Our focus is on Oja's algorithm, which is a one-pass, memory-efficient algorithm requiring only $O(p)$ storage in the rank-one case and $O(pk)$ storage for $k$- PCA. The main goal is to develop a procedure based on Oja's algorithm that is optimal without prior knowledge of the eigengap, which is otherwise needed to tune the learning rates effectively. We do this by first introducing a geometric perspective on the convergence of $k$-PCA: by embedding the Oja iterates into the exterior space, we establish an exact equivalence between $k$-PCA in the ambient space and $1$-PCA in the exterior space. This geometric viewpoint enables us to establish several convergence guarantees, including one based only on a single learning-rate schedule. Building on this geometric perspective, we develop an averaging theory for $k$-PCA and show that the resulting averaged estimator is adaptive: it achieves the nearly optimal convergence rate without prior knowledge of the eigengap. As an application, we construct memory-efficient estimators for elliptical component analysis (ECA) \cite{han2014scale,han2018eca}. Simulation studies and real data analysis are conducted to demonstrate the benefits of our proposed algorithms.

发表机构

  • University of Texas, Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

↑