AI 中文总结
研究针对因果推断中协变量平衡评估问题,扩展柯尔莫哥洛夫-斯米尔诺夫、克莱默-冯·米塞斯和安德森-达林检验以适应加权数据,通过共享程序控制I型错误,保留各检验优势,支持AD为通用默认方法,KS在特定情况有优势。
AI 中文摘要
评估协变量平衡是因果推断中的核心诊断步骤,但常用的汇总度量可能会遗漏其未设计检测的有意义的分布差异。分布拟合优度检验,包括柯尔莫哥洛夫-斯米尔诺夫(KS)、安德森-达林(AD)和克莱默-冯·米塞斯(CVM)检验,提供了更全面的比较,但以前仅适用于未加权数据。我们扩展了这三种检验以适应任何来源的案例权重,使用一种共享的标签置换推断程序,该程序无需假设权重是如何生成的。在一个四场景模拟研究中,在权重变化较大的情况下,从1000到4000的样本量中,所有三种加权检验都将I型错误控制在接近名义水平。对于三种差异类型中的两种,每种检验已知的未加权比较优势在加权下得以保留:KS对中心位置的差异最有效,AD对尾部位置的差异最有效,而AD和CVM对分散差异的表现相当,均优于KS。这些发现支持AD作为常规协变量平衡评估的合理通用默认方法,而当特别怀疑存在中心集中的不平衡时,KS保留优势。这些方法在Stata命令kstest、adtest和cvmtest中实现。
英文摘要
Assessing covariate balance is a core diagnostic step in causal inference, but commonly used summary measures can miss meaningful distributional differences they are not designed to detect. Distributional goodness-of-fit tests, including the Kolmogorov-Smirnov (KS), Anderson-Darling (AD), and Cramer-von Mises (CVM) tests, offer a more complete comparison but have previously been available only for unweighted data. We extend all three to accommodate case weights of any origin, using a shared label-permutation inference procedure that requires no assumption about how the weights were generated. In a four-scenario simulation study, all three weighted tests controlled Type I error close to nominal across sample sizes from 1,000 to 4,000 under substantial weight variability. Each test's known unweighted comparative advantage was preserved under weighting for two of three discrepancy types: KS was most powerful against a centrally located discrepancy, and AD was overwhelmingly most powerful against a tail-located discrepancy, while AD and CVM performed comparably against a diffuse discrepancy, both outperforming KS. These findings support AD as a reasonable general-purpose default for routine covariate balance assessment, while KS retains an advantage when a centrally concentrated imbalance is specifically suspected. The methods are implemented in the Stata commands kstest, adtest, and cvmtest.