AI 中文总结
研究给定混杂因素Z时随机变量X和Y的条件独立检验问题,提出数据自适应分箱扩展LPT,给出任意检验统计量I型错误的有限样本界,证明其功效与最优似然比检验相当,分析线性模型中箱大小影响,表明该策略可行且高效。
AI 中文摘要
在这项工作中,我们研究了在给定混杂因素Z的情况下,检验随机变量X和Y之间条件独立性的问题。局部置换检验(LPT)通过将Z空间划分为预先指定的箱,并在每个箱内对X和Y数据进行置换,来评估观察到的检验统计量的显著性,为该问题提供了一种有原则的方法。然而,当分区预先固定时,结果分区可能不平衡,一些箱可能包含大多数样本,而其他箱只包含少数样本。这促使使用数据自适应分箱策略,如具有固定(通常较小)点数的等大小箱。我们研究了LPT的这种自然且实际重要的扩展,为任意检验统计量的I型错误提供了有限样本界,比先前已知的结果具有更强的有效性。我们还表明,LPT获得的功效与从奈曼 - 皮尔逊引理导出的最优似然比检验相当。在一个线性混杂因素模型类中,我们进一步分析了箱大小的影响,并证明恒定箱大小可以匹配具有不断增长箱大小的分区的性能。这些结果,再加上广泛的数值模拟支持,表明所提出的数据自适应策略既切实可行又具有统计效率。
英文摘要
In this work, we study the problem of testing conditional independence between random variables $X$ and $Y$ given a confounder $Z$. The local permutation test (LPT) offers a principled approach to this problem by partitioning the $Z$-space into pre-specified bins, and permuting the $X$ and $Y$ data within each bin, to assess the significance of an observed test statistic. However, when the partitions are pre-fixed, the resulting partition can be poorly balanced, as some bins may contain most of the samples while others contain only a few. This motivates the use of data-adaptive binning strategies, such as equisized bins with a fixed (typically small) number of points. We study this natural and practically important extension of LPT, providing finite-sample bounds on the Type I error for an arbitrary test statistic, providing stronger validity results than previously known. We also show that LPT attains power comparable to the oracle likelihood ratio tests derived from the Neyman-Pearson lemma. Within a linear confounder model class, we further analyze the effect of bin size and demonstrate that constant bin sizes can match the performance of partitions with growing bin-size. These results, further supported by extensive numerical simulations, position the proposed data-adaptive strategy as both practically implementable and statistically efficient.
Comments88 pages, 7 figures, 2 tables