发表机构
University of Southampton(南安普顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对平行分析在高维混合信号与重尾误差下失效的问题,提出基于秩变换的非参数方法,在宽松条件下一致估计潜在维数,并经模拟与FRED-MD数据验证。
AI 中文摘要
主成分分析和因子分析是高维数据研究的基础,但其有效性完全依赖于正确识别潜在维数 $r$。虽然平行分析(PA)被广泛视为完成此任务的金标准,但在由混合信号和/或重尾误差分布组成的高维情境下,其性能会受到影响。这是因为在上述条件下,应用于数据矩阵 $\ extbf{X}$ 的逐列置换机制(或任何现代变体)无法破坏信号,导致零谱被夸大。为解决此问题,我们提出了一种非参数解决方案。我们证明,潜在维数唯一地保留在秩变换数据矩阵内的有序度中,这恢复了置换下的谱可交换性。我们在宽松条件下证明了秩平行分析(Rank PA)的一致性,既不需要因子的普遍性,也不需要限制性的四阶矩假设。我们确定,Rank PA 能够一致地估计因子数量,最高可达 $O(n^{1/3 - \ au})$ 的最大值,其中 $\ au >0$。我们通过数值模拟和对 FRED-MD 数据集的分析,展示了我们方法的稳健性能。
英文摘要
Principal component analysis and factor analysis are foundational to the study of high-dimensional data, yet their efficacy depends entirely on correctly identifying the latent dimension $r$. While parallel analysis (PA) is widely regarded as the gold standard for this task, its performance suffers in high-dimensional regimes consisting of mixed signals and/or heavy-tailed error distributions. This arises because under the preceding conditions, column-wise permutation mechanism (or any modern variations) applied to the data matrix $\textbf{X}$ fails to destroy the signal, leading to an inflated null spectrum. To address this, we propose a nonparametric solution. We show that the latent dimension is uniquely preserved in the degree of order within the rank-transformed data matrix, which restores spectral exchangeability under permutation. We prove the consistency of Rank PA under relaxed conditions, requiring neither the pervasiveness of factors nor restrictive fourth moment assumptions. We establish that Rank PA can consistently estimate the number of factors up to the maximum of $O(n^{1/3 - τ})$ where $τ>0$. The robust performance of our remedy is demonstrated through numerical simulations and an analysis of the FRED-MD dataset.