AI 中文总结
本文提出鲁棒多任务PCA方法,利用任务间相似性改进特征空间估计并抵抗异常任务,达到极小极大最优速率,通过模拟和真实数据验证有效性。
AI 中文摘要
主成分分析(PCA)是从高维数据中学习低维结构的基本工具。当数据来自多个来源时,潜在的任务分布可能表现出未知程度的相似性,某些任务可能来自任意分布。我们提出了新的多任务PCA程序,利用任务间的相似性结构来改进特征空间估计,同时对异常任务保持鲁棒性。我们建立了非渐近收敛速率,并表明所提出的程序在一系列机制中达到极小极大最优速率。其中一个程序基于Chen、Gao和Ren(2018)的矩阵深度概念,能够实现估计误差对异常任务比例的最优依赖,解决了鲁棒多任务学习中的一个关键挑战。广泛的模拟和真实数据分析证明了所提出方法的有效性。
英文摘要
Principal component analysis (PCA) is a fundamental tool for learning low-dimensional structure from high-dimensional data. When data are collected from multiple sources, the underlying task distributions may exhibit unknown degrees of similarity, with some tasks potentially arising from arbitrary distributions. We propose new multi-task PCA procedures that exploit similarity structure across tasks to improve eigenspace estimation while remaining robust to outlier tasks. We establish non-asymptotic convergence rates and show that the proposed procedures attain minimax optimal rates in a range of regimes. One of the procedures builds on the matrix-depth notion of Chen, Gao, and Ren (2018) and can achieve the optimal dependence of the estimation error on the proportion of outlier tasks, addressing a key challenge in robust multi-task learning. Extensive simulations and real-data analyses demonstrate the effectiveness of the proposed methods.
Comments52 pages, 12 figures. All comments are warmly welcomed