arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10305stat.MEstat.ML

COMPACT:基于完全不可约因果准则的谱调整分数

COMPACT: Spectral Adjustment Scores from a Complete and Irreducible Causal Criterion

Eric V. Strobl

中文总结 AI 辅助

COMPACT算法通过完全不可约因果准则的谱调整分数,解决观测数据因果效应估计的变量选择问题,经模拟和真实数据验证,性能优于多种替代方法。

中文摘要 AI 辅助

观测数据集常包含大量基线变量,但估计因果效应的研究者可能不清楚应将哪些变量纳入调整集,混淆信息也可能分散在多个变量中。倾向评分可通过将高维协变量简化为二元处理的标量来简化调整,尽管倾向评分是最粗糙的平衡分数,但其分布最优性并不意味着对因果图具有最大特异性。我们转而在候选分数、处理和结果中检查所有允许潜在变量的因果图。在忠实性假设下,我们确定最大的无条件和条件依赖关系集合,其真值不受处理是否导致结果的影响,将处理效应估计留给下游分析。该准则定义了可通过这些关系表达的最大特异性图类。随后我们开发了所提出的算法,该算法通过广义特征值问题实现该准则,其分数空间针对平衡坐标和结果引导坐标的张成。我们证明足够信息的代理可在不直接观测调整变量的情况下恢复该张成,刻画所得估计和因果误差,并为完整过程建立自助法有效性。模拟和真实数据应用表明,其性能优于多种替代方法。

英文摘要

Observational datasets frequently contain many baseline variables, yet investigators estimating causal effects may not know which variables to include in the adjustment set. Confounding information may also be distributed weakly across many variables. Propensity scores can simplify adjustment by reducing high-dimensional covariates to a scalar with binary treatment. Although the propensity score is the coarsest balancing score, this distributional optimality does not imply maximal specificity over causal graphs. We instead examine all causal graphs among a candidate score, treatment, and outcome while allowing latent variables. Under faithfulness, we identify the largest set of unconditional and conditional dependence relations whose truth is invariant to whether treatment causes the outcome, leaving treatment-effect estimation to the downstream analysis. This criterion defines the maximally specific graph class expressible through these relations. We then develop the proposed algorithm, which operationalizes the criterion through a generalized eigenvalue problem whose score space targets the span of a balancing coordinate and an outcome-guided coordinate. We show that sufficiently informative proxies can recover this span without direct observation of the adjustment variables, characterize the resulting estimation and causal errors, and establish bootstrap validity for the complete procedure. Simulations and a real-data application demonstrate superior performance over several alternatives.

补充信息

↑