发表机构
School of Population and Public Health, University of British Columbia; Centre for Advancing Health Outcomes, St. Paul’s Hospital(不列颠哥伦比亚大学公共卫生与人口健康学院; 圣保罗医院促进健康成果中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本教程通过真实数据调和 AIPW、TMLE 与 DML 三种估计器,证明共享 nuisance 库和折数后它们结果一致,并强调阳性诊断优先于估计器选择。
AI 中文摘要
增广逆概率加权(AIPW)、目标最大似然估计(TMLE)和双重/去偏机器学习(DML)是通往平均处理效应同一有效影响函数的三种路径——这一成熟理论我们视为背景知识。本教程的贡献在于其在真实数据上基于共享 nuisance 参数的工作化调和:实践者必须使各路径及软件包在哪些方面真正匹配才能达成一致。利用开放的 NHEFS 数据(n=1566)研究戒烟对体重变化的影响,采用一个共享的 Super Learner 库和相同的交叉拟合折数,我们从一个影响函数出发手工构建了全部三种估计器,并给出全样本、交叉拟合和双重交叉拟合三种变体;由此得到的六个双稳健估计值跨度仅为 3.32–3.42 kg,与既有基准一致。随后,我们在自己的引擎与 tmle、AIPW、DoubleML 和 tmle3 软件包之间调和同一估计目标。在默认设置下,这些估计值跨度为 3.32–3.49 kg。一旦库和折数匹配并平均掉单次分割噪声,三个共享库的实现彼此差异在 0.01 kg 以内——因此残余的离散可归因于 nuisance 库、折数和重复次数,而非估计器标签。由于这些估计器共享同一个影响函数,在良好重叠下它们会达成一致;而在存在阳性违背时,若无额外外推假设,合并的 ATE 无法识别,此时它们的有限样本估计可能急剧分化。因此,我们将阳性诊断置于估计器选择之前,在一个无重叠示例上说明失败情形,并以一份报告清单作结。开源 R 代码可复现每一个数字。
英文摘要
Augmented inverse-probability weighting (AIPW), targeted maximum likelihood estimation (TMLE), and double/debiased machine learning (DML) are three routes to the same efficient influence function for the average treatment effect --- settled theory we treat as background. This tutorial's contribution is its worked, shared-nuisance reconciliation on real data: what a practitioner must actually match for the routes, and the software packages, to agree. Working the effect of smoking cessation on weight change in the open NHEFS data (n=1566) with one shared Super Learner library and identical cross-fitting folds, we build all three estimators by hand from one influence function, in full-sample, cross-fit, and double-cross-fit variants; the six resulting doubly-robust estimates span only 3.32--3.42 kg, consistent with the established benchmark. We then reconcile the same estimand across our engine and the tmle, AIPW, DoubleML, and tmle3 packages. At their defaults the estimates span 3.32--3.49 kg. Once the library and folds are matched and single-split noise is averaged out, the three library-sharing implementations agree to within 0.01 kg --- so the residual spread traces to the nuisance library, folds, and repetitions, not to the estimator label. Because the estimators share one influence function, they agree under good overlap; under a positivity violation the pooled ATE is not identified without additional extrapolation assumptions, and their finite-sample estimates can then diverge sharply. We therefore place a positivity diagnosis ahead of estimator choice, illustrate the failure on a no-overlap example, and close with a reporting checklist. Open-source R code reproduces every number.
Comments45 pages, 8 figures, 14 tables. Tutorial. R code at https://github.com/ehsanx/dml-alternatives, R package at https://github.com/ehsanx/dr3/