发表机构
Universität Potsdam; Cornell University(波茨坦大学; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出通过施加层级秩结构来正则化小样本协方差矩阵估计,相比稀疏性方法适用更广,并给出高效算法,实验证明其估计误差更小。
AI 中文摘要
我们考虑从非常有限的样本数量中估计高维协方差矩阵的问题。这个问题在计算流体力学中普遍存在,其中必须使用少量流体快照来构造确定降阶模型的格拉姆矩阵;在计算地球科学中也是如此,其中必须使用少量地球系统预报集合来估计与预报不确定性相关的协方差矩阵。通常的做法是通过施加“局地化”结构来正则化小样本协方差,该结构强制实施物理上合理的相关长度尺度,施加稀疏性约束,向指定目标“收缩”,或衰减小相关性。我们提出了一种替代技术,通过施加层级秩结构来正则化小样本协方差。与假设稀疏性的正则化方法(如空间局地化)相比,层级秩结构能够容纳更广泛的协方差矩阵,大致对应于长程相关性比短程相关性变化更平滑的情况。它还产生一种数据稀疏的矩阵格式,允许高效的矩阵-向量乘积。我们提出了理论和算法,展示了如何从有限样本中高效估计高维层级秩结构协方差矩阵。通过误差分析和各种模型问题的数值实验,我们证明了这些技术在减少采样误差方面是有效的,并且在许多情况下,它们比传统技术实现了更小的估计误差。
英文摘要
We consider the problem of estimating a high-dimensional covariance matrix from a very limited number of samples. This problem is ubiquitous in computational fluid dynamics, where a small number of fluid snapshots must be used to construct a Gramian matrix determining a reduced-order model, as well as in computational geoscience, where a small ensemble of Earth system forecasts must be used to estimate the covariance matrix associated with the forecast uncertainty. It is common practice to regularize the small-sample covariance by imposing a "localization" structure that enforces a physically realistic correlation length scale, imposing a sparsity constraint, "shrinking" towards a prescribed target, or attenuating small correlations. We propose an alternate technique that regularizes the small-sample covariance by imposing hierarchical rank structure. Compared to regularization methods that assume sparsity such as spatial localization, hierarchical rank structure accommodates a wider range of covariance matrices, roughly corresponding to situations where long-range correlations vary more smoothly than short-range ones. It also results in a data-sparse matrix format that permits highly efficient matrix-vector products. We present theory and algorithms which show how to efficiently estimate a high-dimensional, hierarchically rank structured covariance matrix from limited samples. Through an error analysis and numerical experiments with a variety of model problems, we demonstrate that these techniques are effective at reducing sampling errors, and that in many cases they achieve smaller estimation error than conventional techniques.