未测量混杂下因果效应估计的近端平衡
Proximal Balancing for Causal Effect Estimation under Unmeasured Confounding
AI总结:
针对未测量混杂下的因果效应估计,提出近端平衡方法,通过学习协变量和代理的低维摘要实现组间可比,无需逆问题或潜模型,并提供理论保证与PROBE算法。
AI中文摘要:
从观测数据中估计因果效应对科学和政策至关重要,但当混杂因素未被测量时,效应无法被识别。近端因果推断利用未测量混杂因素的代理变量来解决此问题。然而,现有的基于代理的方法要么指定代理角色并求解逆问题,该问题不适定且在高维代理下难以估计,要么使用潜变量模型,该模型假设学习到的潜变量与隐藏混杂因素匹配,若不匹配则留下偏差。为应对这些挑战,我们引入了近端平衡。它将协变量平衡的经典思想推广到仅通过代理观测的混杂因素:它学习协变量和代理的低维摘要,使处理组具有可比性,然后针对该摘要进行调整。它不需要指定的代理角色、逆问题或潜模型。我们给出了识别理论、有限样本保证以及实用算法PROBE。我们在低维、高维和图像代理以及真实世界数据上展示了该方法。
英文摘要:
Estimating causal effects from observational data is central to science and policy, but the effects are not identified when confounders are unmeasured. Proximal causal inference addresses this problem with proxies of the unmeasured confounders. However, existing proxy-based approaches either designate proxy roles and solve an inverse problem, which is ill-posed and hard to estimate with high-dimensional proxies, or use a latent-variable model, which assumes that the learned latent variable matches the hidden confounder and leaves bias when it does not. To address these challenges, we introduce proximal balancing. It carries the classical idea of covariate balancing to confounders that are observed only through proxies: it learns a low-dimensional summary of the covariates and proxies that makes the treatment groups comparable, and then adjusts for this summary. It needs no designated proxy roles, inverse problem, or latent model. We give identification theory, finite-sample guarantees, and a practical algorithm, PROBE. We demonstrate the method on low-dimensional, high-dimensional, and image proxies and on real-world data.