arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09234stat.ME

自删失模型下信息性缺失的纵向因果推断

Longitudinal causal inference under informative missingness using the self-censoring model

Iván Díaz

首次发表
浏览论文内容

中文总结 AI 辅助

针对纵向数据信息性缺失问题,提出自删失模型下的因果识别与估计方法,结合影子变量条件矩限制实现反事实均值递归识别,并构建渐近线性估计器以支持机器学习与根n推断。

中文摘要 AI 辅助

纵向观测数据常常受到信息性缺失的影响。例如,在使用电子健康记录的医疗保健研究中,协变量和结局仅在患者与医疗系统互动时才被测量。这些数据是否被观测通常取决于未观测到的变量值本身,从而引入了一种常见的数据非随机缺失形式。我们开发了在自删失模型下纵向修正治疗策略效应的识别与估计方法,该模型明确建模了这种依赖性。我们考虑仅在测量时间(例如与医疗系统的互动)可以修改的治疗效应。我们研究了两种设置:一种设置中观测历史足以控制治疗混杂,另一种设置中混杂控制需要未在错过的访视中测量的协变量值。在两种设置中,我们将纵向因果识别与影子变量条件矩限制相结合,以获得反事实结局均值的递归识别公式。我们推导了路径导数梯度和二阶余项展开,并利用这些结果构建了能够适应机器学习进行干扰函数估计的估计器。在适当的产品率条件和交叉拟合下,这些估计器是渐近线性的,并允许根n推断,包括跨结局时间的同时推断。我们还描述了用于估计我们方法所需干扰函数的条件矩学习和集成方法,以及一种在稀疏测量的结局中借用信息的池化和平滑策略。

英文摘要

Longitudinal observational data are often subject to informative missingness. For example, in healthcare research using electronic health records, covariates and outcomes are measured only when patients interact with the healthcare system. Whether these data are observed usually depends on the unobserved variable values themselves, introducing a common form of data missing not at random. We develop identification and estimation methods for the effects of longitudinal modified treatment policies under a self-censoring model which explicitly models this dependence. We consider effects of treatments that can be modified only at measurement times (e.g., interactions with the heathcare system). We study two settings: one in which the observed history suffices for treatment confounding control, and another in which confounder control requires covariate values that were not measured at missed visits. In both settings, we combine longitudinal causal identification with shadow-variable conditional moment restrictions to obtain recursive identification formulas for the counterfactual outcome mean. We derive gradients of the pathwise derivatives and second-order remainder expansions, and use these results to construct estimators that accommodate machine learning for nuisance-function estimation. Under suitable product-rate conditions and cross-fitting, the estimators are asymptotically linear and permit root-n inference, including simultaneous inference across outcome times. We also describe conditional-moment learning and ensemble approaches for estimating the nuisance functions required by our approach, together with a pool-and-smooth strategy that borrows information across sparsely measured outcomes.

发表机构

  • New York University Grossman School of Medicine(纽约大学格罗斯曼医学院)

机构由 AI 辅助整理,请以论文原文为准。

↑