持久图统计推断的界标嵌入方法:极小极大理论与有限逼近
Statistical Inference for Persistence Diagrams via Landmark Embeddings: Minimax Theory and Finite Approximation
AI总结:
针对持久图总体推断,提出基于加性界标嵌入(PLACE/PALACE)的框架,通过下失真证书和有限样本理论实现均值分离下界与置信推断,并经模拟和脑影像数据验证。
AI中文摘要:
希尔伯特空间嵌入使得对持久图总体的推断成为可能,但单个图之间的分离在总体平均后未必能保持。我们开发了一个针对总体均值嵌入的推断框架,特别关注加性界标表示PLACE和PALACE。将每个图视为一个独立观测,我们应用希尔伯特空间极限理论,在适当的矩条件下获得协方差估计量、双样本检验和置信球,且无需下失真界。对于加性嵌入,我们将总体均值识别为均值计数测度的嵌入,并表明仅凭这些测度的几何分离无法保证一致的检验功效。随后,我们引入一个包含潜在模板图、缺失特征和位置扰动的模型。在共同或特征特定的流行条件下,图级下失真证书为总体均值分离提供了显式下界。这些间隔提供了有限样本的一致功效保证,额外的信息散度比较在受限尺度范围内给出了匹配的样本复杂度界。置信集给出了总体均值测度传输分离的下界,以及对指定结构化备择的排除保证。我们还量化了正交截断如何改变认证信号以及针对完整嵌入的置信集所需的逼近余量,将样本量、保留坐标和模板分离联系起来。模拟实验检验了校准、功效和覆盖率,对自闭症脑影像数据交换数据库中静息态连接性的分析展示了这些程序的应用。
英文摘要:
Hilbert-space embeddings enable inference for populations of persistence diagrams, but separation between individual diagrams need not survive population averaging. We develop a framework for inference on population mean embeddings, with particular attention to the additive landmark representations PLACE and PALACE. Treating each diagram as one independent observation, we apply Hilbert-space limit theory to obtain covariance estimators, two-sample tests, and confidence balls under suitable moment conditions, without requiring a lower-distortion bound. For additive embeddings, we identify the population mean as an embedding of the mean counting measure and show that geometric separation of these measures alone cannot guarantee uniform testing power. We then introduce a model with latent template diagrams, missing features, and location perturbations. Under common or feature-specific prevalence conditions, a diagram-level lower-distortion certificate yields explicit lower bounds on population mean separation. These margins provide finite-sample uniform power guarantees, and an additional information-divergence comparison gives matching sample-complexity bounds over restricted scale ranges. Confidence sets yield lower bounds on transport separation of population mean measures and exclusion guarantees for specified structured alternatives. We also quantify how orthogonal truncation changes the certified signal and the approximation allowance needed for confidence sets targeting the full embedding, relating sample size, retained coordinates, and template separation. Simulations examine calibration, power, and coverage, and an analysis of resting-state connectivity from the Autism Brain Imaging Data Exchange illustrates the procedures.