网络干扰下通过邻域排除交叉拟合的无偏处理效应估计
Unbiased Treatment Effect Estimation under Network Interference via Neighborhood-Excluded Cross-Fitting
- Tsinghua University(清华大学)
- Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对网络干扰下的处理效应估计,提出邻域排除交叉拟合方法,恢复有限样本无偏性,并给出渐近有效推断及最优调整程序,模拟和实例验证其精度优势。
AI中文摘要:
在没有干扰的情况下,交叉拟合能够在独立单元级随机化下进行灵活的协变量调整,同时保持有限样本无偏性。在网络干扰下,仅靠样本外预测不再保证无偏性:进入评估折霍维茨-汤普森权重的分配也可能影响训练样本中的结果,导致拟合预测与这些权重之间产生依赖性。我们开发了邻域排除交叉拟合方法,该方法构建针对估计目标和设计特定的训练样本,以恢复有限样本无偏性所需的条件独立性,而无需正确指定的结果模型。我们在伯努利随机化下为直接效应和间接效应建立了渐近有效的基于设计的沃尔德推断,并在伯努利整群随机化下为全局平均处理效应建立了同样的推断。邻域排除在选择折数时产生权衡:单元级分割可能要求折数随平均排除邻域大小增长,而整群级分割可以大幅放宽这一要求,在部分干扰下允许固定折数。对于伯努利随机化下的线性调整,我们推导了方差最优和置信区间长度最优的程序,建立了允许协变量维度发散的显式速率条件,并证明了方差最优程序渐近无害。模拟实验展示了省略邻域排除带来的偏差以及调整带来的精度提升。一项应用于社交网络实验的实例得到的置信区间比未调整估计器的置信区间大幅缩短。
英文摘要:
Without interference, cross-fitting enables flexible covariate adjustment while preserving finite-sample unbiasedness under independent unit-level randomization. Under network interference, out-of-sample prediction alone no longer guarantees unbiasedness: assignments entering evaluation-fold Horvitz--Thompson weights may also affect outcomes in the training sample, inducing dependence between fitted predictions and those weights. We develop neighborhood-excluded cross-fitting, which constructs estimand- and design-specific training samples to restore the conditional independence needed for finite-sample unbiasedness without a correctly specified outcome model. We establish asymptotically valid design-based Wald inference for direct and indirect effects under Bernoulli randomization and for the global average treatment effect under Bernoulli cluster randomization. Neighborhood exclusion creates a trade-off in choosing the number of folds: unit-level splitting may require the number of folds to grow with average exclusion-neighborhood size, while cluster-level splitting can substantially relaxes this requirement, permitting a fixed number of folds under partial interference. For linear adjustment under Bernoulli randomization, we derive variance-optimal and confidence-interval-length-optimal procedures, establish explicit rate conditions allowing the covariate dimension to diverge, and show that the variance-optimal procedure is asymptotically no-harm. Simulations illustrate the bias from omitting neighborhood exclusion and the precision gains from adjustment. An application to a social network experiment yields confidence intervals substantially shorter than those from the unadjusted estimator.