发表机构
The University of Tokyo; Hasso Plattner Institute for Digital Engineering; Institute of Statistical Mathematics; The Graduate University for Advanced Studies, SOKENDAI; RIKEN Centre for Advanced Intelligence Project (AIP); Hasso Plattner Institute for Digital Health at the Icahn School of Medicine at Mount Sinai(东京大学; 哈索·普拉特纳数字工程研究所; 统计数学研究所; 综合研究大学院大学(SOKENDAI); 理化学研究所先进智能项目中心(AIP); 西奈山伊坎医学院哈索·普拉特纳数字健康研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出利用随机插值器连接两分布至高斯瓶颈,通过路径对称性检验双样本是否同分布,聚合去噪器与速度场的回归风险,所得检验在有限样本有效且功效显著提升。
AI 中文摘要
现代深度生成模型主要因其生成逼真样本的能力而被研究,然而它们学到的生成动力学也可以作为统计推断的对象。我们将这一思想应用于双样本检验,即判断两个有限数据集是否由同一分布生成的问题。利用随机插值器,我们将两个分布连接到一个共享的高斯瓶颈,使得所得路径的每一半都是作用于单一总体的高斯信道。我们证明,当且仅当总体去噪器(或等价地,两半的速度场)在任意单一噪声水平上重合时,原假设成立,这相当于路径关于瓶颈的反射对称性。偏离这种对称性会产生连续的双样本见证者,我们通过在学习到的去噪器和速度场上计算的留出回归风险来估计它们,并沿路径进行聚合;在信息论加权下,聚合的差异等于噪声平滑分布之间的杰弗里斯散度。通过置换校准所得统计量,得到的检验对于任何训练好的网络在有限样本中都是有效的,并且当场被准确学习时具有一致性。在一个合成基准和三个图像基准上,所提出的检验在相等的总样本预算下,将功效相对于最强基线提高了最多33个百分点,最佳回归表示和路径加权的选择取决于数据模态。这些结果表明,生成路径为统计检验提供了原则性的表示,将随机插值模型扩展到了生成之外。
英文摘要
Modern deep generative models are primarily studied for their ability to generate realistic samples, yet the generative dynamics they learn can also serve as objects of statistical inference. We develop this idea for two-sample testing, the problem of deciding whether the same distribution generated two finite datasets. Using stochastic interpolants, we connect both distributions to a shared Gaussian bottleneck, so that each half of the resulting path is a Gaussian channel acting on a single population. We prove that the null hypothesis holds if and only if the population denoiser, or equivalently, the velocity fields of the two halves, coincide at any single noise level, which amounts to a reflection symmetry of the path about the bottleneck. Deviations from this symmetry yield a continuum of two-sample witnesses, which we estimate via held-out regression risks on learned denoisers and velocities and aggregate along the path; under an information-theoretic weighting, the aggregated discrepancy equals the Jeffreys divergence between the noise-smoothed distributions. Calibrating the resulting statistics by permutation yields tests that are valid in finite samples for any trained networks and consistent when the fields are learned accurately. On a synthetic benchmark and three image benchmarks, the proposed tests improve power over the strongest baseline by up to 33 percentage points at an equal total sample budget, with the best choice of regression representation and path weighting depending on the data modality. These results show that generative paths provide a principled representation for statistical testing, extending stochastic-interpolant models beyond generation.