arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于旋转的子空间跟踪用于流式数据上的鲁棒核主成分分析

Rotation-Based Subspace Tracking for Robust Kernel PCA on Streaming Data

Kris Lokere, John Fossaceca

arXiv 2609.15488首次发表:更新:

发表机构

Harvard University; George Washington University(哈佛大学; 乔治华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于旋转的更新机制,在再生核希尔伯特空间中向新样本旋转子空间估计,结合鲁棒影响函数抑制异常值,实现流式数据上核PCA的动态跟踪,实验证明比梯度下降收敛更快。

AI 中文摘要

机器学习模型处理大量数据,主成分分析(PCA)是一种广泛使用的技术,用于降低数据维度并提取有用特征。在实践中,数据集经常随时间变化(数据漂移)和/或一次到达一个样本(流式数据),这使得在批量模式下一次性处理整个数据集变得不可行。真实世界的数据也经常包含非线性模式,传统PCA无法提取这些模式。核PCA通过将样本隐式映射到再生核希尔伯特空间(RKHS)来解决这一问题。原始数据也经常包含异常值,除非算法具有鲁棒性,否则这些异常值可能对估计的子空间产生过大的影响。然而,现有的在线鲁棒核PCA算法被设计为收敛到假设固定的子空间,一旦实现这种初始对齐,基于梯度下降的更新在跟踪进一步变化时失去效力。本文介绍了一种基于旋转的更新机制,该机制通过在再生核希尔伯特空间中将子空间估计向每个新到来的特征向量旋转来更新子空间估计,而不是仅依赖梯度下降。我们提出了两种互补的旋转策略,并表明旋转的程度可以通过鲁棒影响函数进行调节,以减轻异常值的影响。通过在具有已知真实子空间的合成流式数据上的实验,我们表明逐样本旋转比单独使用梯度下降收敛更快,展示了在流式数据中动态跟踪非线性子空间的有效机制。

英文摘要

Machine learning models process large amounts of data, and Principal Component Analysis (PCA) is a widely used technique to reduce the dimensionality of the data and extract useful features. In practice, datasets often change over time (data drift) and/or arrive one sample at a time (streaming data), making it infeasible to process the entire dataset at once in batch mode. Real-world data also often contains nonlinear patterns, which traditional PCA cannot extract. Kernel PCA addresses this by implicitly mapping samples into a Reproducing Kernel Hilbert Space (RKHS). Raw data also often contains outliers, which can have an outsized effect on the estimated subspace unless the algorithm is made robust. However, existing online robust kernel PCA algorithms are designed to converge to a subspace that is assumed to be fixed, and gradient-descent-based updates lose their effectiveness at tracking further changes once this initial alignment is achieved. This paper introduces a rotation-based update mechanism, which updates the subspace estimate by rotating it toward each new incoming feature vector in Reproducing Kernel Hilbert Space, rather than relying on gradient descent alone. We present two complementary rotation strategies, and show that the extent of rotation can be moderated by a robust influence function to mitigate the effect of outliers. Through experiments on synthetic streaming data with a known ground-truth subspace, we show that per-sample rotations converge faster than gradient descent alone, demonstrating an effective mechanism for dynamically tracking a nonlinear subspace in streaming data.

Comments7 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑