蒙特卡罗检验:无需联合可交换性的非渐近保证
Monte Carlo testing: non-asymptotic guarantees without joint exchangeability
浏览论文内容
中文总结 AI 辅助
研究假设检验中蒙特卡罗检验问题,针对常见计算程序不足,给出有限样本保证。条件是观测数据和副本在给定潜在变量下条件独立同分布,对任意统计量和副本数成立,解释有限样本振荡行为,还为其他推断问题提供见解。
中文摘要 AI 辅助
在假设检验中,蒙特卡罗检验通常通过精确的原假设模拟或观测数据及其模拟副本的联合可交换性来证明其合理性。这为常见的计算程序留下了空白,例如在观测数据处初始化的并行MCMC采样,其中每个副本可能在边际上服从原假设,甚至与观测值成对可交换,但整个集合不是联合可交换的。在这种情况下,当链未混合时,通常的经验p值可能无效,而像Besag-Clifford中心辐条采样器这样的精确可交换构造可能会受到高条件蒙特卡罗变异性的影响。我们给出了这种中间情况的有限样本保证。如果在原假设下,观测数据$X$和副本$X'\sim P(\cdot\mid X)$在给定一个潜在变量的条件下是独立同分布的,那么对于任何预先指定的统计量和任何有限数量$m$的条件独立蒙特卡罗副本,所得的经验p值满足$\mathbb P\{p_m\le \alpha\}\le 2\alpha$。这个保证不需要混合条件,并且对任何副本数量$m$都成立,它解释了通过MCMC采样进行推断时的有限样本振荡行为。此外,我们进一步表明,该保证为其他设置中出现的推断问题提供了见解,包括对贝叶斯模型的推断(恢复了一个经典结果,表明后验预测p值的有效性高达2倍),以及通过平衡置换检验进行的推断。
英文摘要
In hypothesis testing, Monte Carlo tests are usually justified either by exact null simulation or by joint exchangeability of the observed data and its simulated copies. This leaves a gap for common computational procedures, such as parallel MCMC sampling initialized at the observed data, where each copy may be marginally null and even pairwise exchangeable with the observation, but the full collection is not jointly exchangeable. In such cases the usual empirical p-value can be invalid when the chain has not mixed, while exactly exchangeable constructions such as the Besag--Clifford hub-and-spoke sampler may suffer from high conditional Monte Carlo variability. We give finite-sample guarantees for this intermediate regime. If, under the null, the observed data $X$ and a copy $X'\sim P(\cdot\mid X)$ are conditionally i.i.d.\ given a latent variable, then for any prespecified statistic and any finite number $m$ of conditionally independent Monte Carlo copies, the resulting empirical p-value obeys $\mathbb P\{p_m\le α\}\le 2α$. This guarantee requires no mixing conditions and holds for any number of copies $m$, and it explains finite-sample oscillatory behavior in inference via MCMC sampling. In addition, we further show that the guarantee provides insights into inference problems arising in other settings, including inference on Bayesian models (recovering a classical result showing validity up to a factor of $2$ for posterior predictive p-values), and inference via balanced permutation tests.