发表机构
University of Dhaka(达卡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对非随机缺失数据下因果估计量不可识别的问题,提出基于完备性条件的识别框架和EM估计算法,并扩展至中介分析;模拟显示偏差显著低于完整案例分析与多重插补,NHANES应用验证了教育对抑郁的因果效应。
AI 中文摘要
非随机缺失(MNAR)数据对因果推断构成重大挑战,尤其是在混杂因素和结局变量均部分缺失的情况下。若不在因果推断所需假设之外增加额外假设,因果估计量在MNAR机制下通常不可识别。本文首先在若干合理的MNAR机制下,利用完备性条件建立了因果估计量——平均处理效应——的识别结果,并提出了一种基于期望最大化(EM)算法的估计方法。我们进一步将该识别与估计框架扩展到中介分析,使得在MNAR机制下能够估计自然直接效应和自然间接效应。通过大量模拟研究,我们将所提方法与两种广泛使用的缺失数据处理方法——完整案例分析法和多重插补法——进行比较。结果表明,在所考虑的MNAR机制下,所提方法产生的偏差显著更低。最后,我们将所提方法应用于NHANES数据,以估计教育对抑郁的因果效应,其中健康状况作为中介变量。
英文摘要
Missing not at random (MNAR) data pose significant challenges for causal inference, particularly when both confounders and the outcome are partially observed. Without additional assumptions beyond those required for causal inference, causal estimands are generally not identifiable under MNAR mechanisms. This paper first develops identification results for the causal estimand, the average treatment effect, under several plausible MNAR mechanisms using completeness conditions, and proposes an estimation approach based on the Expectation-Maximization (EM) algorithm. We further extend this identification and estimation framework to mediation analysis, enabling the estimation of natural direct and indirect effects under MNAR mechanisms. Through extensive simulation studies, we compare the proposed method with two widely used approaches for handling missing data, complete-case analysis and multiple imputation. The results show that the proposed method yields substantially lower bias under the considered MNAR mechanisms. Finally, we apply the proposed approach to NHANES data to estimate the causal effect of education on depression, with health condition as the mediator.
Comments26 pages, 6 figures, and 5 tables