arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

比较在时变混杂因素下估计平均治疗效果的缺失数据方法:一项模拟研究

Comparing Missing Data Methods for Estimating Average Treatment Effects Under Time-Varying Confounding: A Simulation Study

Ben Swallow, Lars Brestrich, Victor Velasco-Pardo

arXiv 2607.17775首次发表:更新:

AI 中文总结

研究在时变混杂因素下估计平均治疗效果的缺失数据方法,通过模拟比较分层热卡插补、单模插补、MICE和完整病例分析等方法,发现性能受缺失机制等因素影响,多重插补表现较好,为相关研究提供参考。

AI 中文摘要

在实际统计应用中,缺失数据和混杂因素很常见。但很少有研究探讨二元变量在时变混杂因素下插补方法的表现,以及缺失机制、缺失率、缺失位置和样本量如何共同影响性能和潜在的可识别性条件。我们生成了合成数据并进行模拟研究,比较不同因素场景下的缺失数据方法。在处理和结果变量中引入缺失值,应用分层热卡插补、单模插补、链式方程多重插补(MICE)和完整病例分析。使用倾向得分加权的逻辑回归获得平均治疗效果(ATE)估计值,并在48种场景下各重复500次测量覆盖率、绝对偏差和经验标准误差。性能主要由缺失机制和方法选择驱动,多重插补通常比其他方法具有更好的覆盖率和更低的偏差。缺失位置也很重要,而缺失率和样本量主要影响正性违背,在非随机缺失、高缺失率和小样本量下最为明显。倾向得分模型能充分控制从轻度到中度混杂的可交换性违背,而强混杂会使覆盖率略有下降。进一步的研究应使用更先进的方法和更复杂的缺失场景,探讨在缺失情况下可识别性条件可能被违反的其他方式。

英文摘要

Missing data and confounding are common in real-world statistical applications, yet few studies have examined how imputation methods perform under time-varying confounding in binary variables, or how missingness mechanism, missing rate, missingness location and sample size jointly affect performance and the underlying identifiability conditions. We generated synthetic data and conducted a simulation study comparing missing data methods across scenarios varying these factors. Missingness was introduced in both treatment and outcome variables, and we applied stratified hot deck imputation, single mode imputation, multiple imputation with chained equations (MICE), and complete-case analysis. Average treatment effect (ATE) estimates were obtained using logistic regression with propensity score weighting, and we measured coverage, absolute bias and empirical standard errors across 48 scenarios with 500 replications each. Performance was primarily driven by the missingness mechanism and choice of method, with multiple imputation generally achieving better coverage and lower bias than other methods. Missingness location was also important, while missing rate and sample size primarily affected positivity violations, which were most pronounced under MNAR, high missingness and low sample sizes. Exchangeability violations from mild to moderate confounding were adequately controlled for by propensity score models, whereas strong confounding produced a modest decrease in coverage. Further research should examine additional ways identifiability conditions can be violated under missingness, using more advanced methods and more complex missingness scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑