用于函数数据两样本检验的混合高斯投影
Mixed Gaussian Projections for Two-Sample Testing of Functional Data
浏览论文内容
中文总结 AI 辅助
本文提出结合Haar与Fourier高斯分量的混合高斯投影方法,用于函数数据两样本检验,可提升检验功效,在ECG5000数据集上识别出两类间的局部化差异。
中文摘要 AI 辅助
函数数据的随机投影检验依赖于生成投影方向的概率法则,需选择合适的测度来生成数据投影的随机方向。在L²空间中,可利用高斯测度区分由矩定义的概率测度;高斯投影法则的协方差算子决定了函数空间中哪些区域和结构会获得可观概率,进而影响有限样本功效。本文表明,非退化高斯测度的混合保持了几乎必然分离性,并在函数概率测度间诱导出一种度量;集中论证进一步明确,投影检验的功效由所选法则关联的积分投影距离决定。作为具体构造,本文结合了Haar与Fourier高斯分量,二者分别强调局部化与振荡性偏差;所得置换检验保留了一致性,且分量标签与上尾分离得分可对检测到的差异几何提供描述性指示。模拟结果显示了两个分量的专门化特性及其混合的鲁棒性;ECG5000应用识别出两类间主要为局部化差异。本文还与合并函数主成分分析协方差算子进行了比较,该有限秩、数据自适应的投影几何因具有标签不变构造而保持置换有效性,并在多种模拟设置(尤其协方差结构变化下)提升了功效。
英文摘要
Random-projection tests for functional data depend on the probability law used to generate projection directions. A measure has to be selected to generate the random directions in which the data is projected. In $L^2$, probability measures defined by their moments can be discriminated using Gaussian measures. The covariance operator of the Gaussian projection law determines which regions and structures of the functional space receive appreciable probability, and consequently affects finite-sample power. We show that mixtures of non-degenerate Gaussian measures preserve the almost-sure separation property and induce a metric between functional probability laws. A concentration argument further makes explicit that the power of projection tests is governed by the integrated projected distance associated with the chosen law. As a concrete construction, we combine Haar and Fourier Gaussian components, which emphasize localized and oscillatory departures, respectively. The resulting permutation test retains consistency, while component labels and upper-tail separation scores provide a descriptive indication of the geometry of the detected discrepancy. Simulations illustrate the specialization of the two components and the robustness of their mixture, and an ECG5000 application identifies a predominantly localized difference between two classes. We also compare with a pooled functional principal component analysis covariance operator. This finite-rank, data-adaptive projection geometry preserves permutation validity through its label-invariant construction and improves power in several simulated settings, particularly under a change in covariance structure.