AI 中文总结
本文针对高维因子模型中主成分估计误差,将其分解为子空间外误差和子空间内误差两项,并证明渐近性质,通过三因子模拟验证,为误差量化提供方法。
AI 中文摘要
在统计因子模型中,样本协方差矩阵的主成分(或特征向量)被用作{\it 主方向}的估计,主方向是观测变量集合共同运动的真实驱动因素。我们将这些估计中通常显著的误差写成两个可解释项之和,并证明当变量数量增长而样本量有界时,这两项具有几乎必然的渐近极限。这种情况在金融经济学、基因组学、机器学习和信号处理中很常见。{\it 子空间外误差}衡量估计到总体因子暴露所张成子空间的距离。它可以用数据表示,从而提供了一个可估计的误差下限。{\it 子空间内误差}源于潜在因子回报的固定样本量,无法仅从数据估计。我们通过美国公开股票市场的三因子模拟来说明我们的误差分析,展示了误差及其组成部分的大小对维度和样本量的依赖性。在该模拟中,子空间外误差占主导地位。依赖主成分分析来估计因子模型的研究人员可以使用我们的结果来量化基于模型的预测和归因中的误差。
英文摘要
In a statistical factor model, principal components (or eigenvectors) of a sample covariance matrix serve as estimates of {\it principal directions}, the true drivers of co-movement of a collection of observed variables. We write the often substantial error in these estimates as a sum of two interpretable terms, which we show have almost sure asymptotic limits as the number of variables grows with sample size bounded. This scenario is commonplace in financial economics, genomics, machine learning and signal processing. {\it Out-of-subspace error} measures the distance from an estimate to the subspace spanned by population factor exposures. It can be expressed in terms of data, providing an estimable floor for error. {\it In-subspace error} arises from the fixed sample size of the latent factor returns and cannot be estimated from data alone. We illustrate our error analysis with a three-factor simulation of the US public equity market, showing the dependence of the magnitude of the error and its components on dimension and sample size. In that simulation, out-of-subspace error dominates. Researchers who rely on principal component analysis to estimate factor models can use our results to quantify errors in model-based predictions and attributions.
Comments35 complied pages, 4 figures