可分块累积量张量:最优估计及其在非高斯数据建模中的应用
Bandable Cumulant Tensors: Optimal Estimation and Applications in Non-Gaussian Data Modeling
浏览论文内容
中文总结 AI 辅助
该研究针对高维非高斯数据,提出可分块累积量类与锥形样本累积量估计器,证明其最优性并将其应用于多种下游任务,经模拟和真实数据验证有效。
中文摘要 AI 辅助
高阶累积量能够捕获协方差所遗漏的非高斯相关性,但在高维场景中难以应用。d阶累积量张量有p^d个条目,而插件式样本累积量在张量谱范数下通常甚至无法达到速率最优。对于有序数据,两个难题可通过一种补救方案解决:假设高阶交互作用沿张量主对角线衰减,我们引入可分块累积量类和锥形样本累积量估计器,该估计器仅计算带宽k处的O(pk^{d-1})个局部条目,从不构建完整张量。在指数型尾条件下,我们证明了非渐近谱范数界,将锥形偏差与随机误差分离,并在同一类上得到极小极大下界,该下界与上界的主导偏差和随机项匹配;对于次高斯观测,当n≫(k+log p)^{d-1}时,锥形估计器在先知带宽k处达到极小极大速率,环境维度仅通过log p进入。定位还抑制了导致次优性的高阶波动,因此锥形估计在此处的作用强于可分块协方差估计。谱范数保证直接传递到下游任务,为自回归模型中的累积量Yule–Walker估计、移动平均模型中的最小距离估计以及传感器阵列中的匹配滤波器源定位提供插件误差界。模拟验证了该理论,对RR间期、空气质量和Neuropixels记录的真实数据分析表明了由此产生的稳定性提升。
英文摘要
Higher-order cumulants capture the non-Gaussian dependence that covariance misses, but they are hard to use in high dimensions. An order-$d$ cumulant tensor has $p^d$ entries, and the plug-in sample cumulant is generally not even rate-optimal under the tensor spectral norm. For ordered data, both difficulties admit one remedy: assuming that higher-order interactions decay away from the main tensor diagonal, we introduce a bandable cumulant class and a tapered sample cumulant estimator that computes only $O(pk^{d-1})$ local entries at bandwidth $k$ and never forms the full tensor. Under exponential-type tail conditions, we prove nonasymptotic spectral-norm bounds that separate tapering bias from stochastic error, and a minimax lower bound over the same class that matches the leading bias and stochastic terms of the upper bound; for sub-Gaussian observations, the tapered estimator attains the minimax rate whenever $n\gg(k+\log p)^{d-1}$ at the oracle bandwidth $k$, with the ambient dimension entering only through $\log p$. Localization also suppresses the higher-order fluctuations behind this suboptimality, so tapering plays a stronger role here than in bandable covariance estimation. The spectral-norm guarantee transfers directly to downstream tasks, yielding plug-in error bounds for cumulant Yule--Walker estimation in autoregressive models, minimum-distance estimation in moving-average models, and matched-filter source localization in sensor arrays. Simulations corroborate the theory, and real-data analyses of RR-interval, air-quality, and Neuropixels recordings illustrate the resulting stability gains.