AI 中文总结
该研究构建时变处理的基准,发现纵向匹配与靶试验模拟方法针对不同因果估计量,量化幻影偏倚、存在排名反转、方差估计有差异、模型敏感性受校准影响,部分结果可在另一机制重复。
AI 中文摘要
在不可折叠的生存机制下,纵向匹配与靶试验模拟方法并非针对同一真值的竞争估计量,而是不同因果问题的解答,因此以单一“真实风险比”对二者评分的基准会产生偏倚。我们提供了相对效率、方差估计与模型敏感性的直接比较,弥补了现有研究的不足。在具有已知真值的特意构造的不可折叠连续时间Cox机制中,主流方法族(序贯Cox、序贯分层、风险集匹配、处理加权逆概率(IPTW)边际结构模型)针对数值上不同的因果估计量(边际、条件、两种处理组平均处理效应、意向治疗与符合方案效应)。第一,我们量化了共享边际真值产生的幻影偏倚:匹配估计量为0.32-0.33对数累积风险比单位,条件方法为0.15;关联型 naive 时变Cox偏倚达0.76,该总差异因混杂与估计量缺口而放大。第二,存在排名反转:推荐方法随目标估计量变化,低方差的非目标估计量仍可在均方误差上胜出。第三,跨族方差结果:聚类稳健三明治估计对试验堆叠估计量的覆盖率接近名义值(0.90),但对匹配估计量覆盖率不足(0.77-0.82),预先设定的n=500自助法子研究可将其提升至0.95-0.96。第四,模型敏感性:遗漏一个混杂因素会导致0.45-0.50对数风险比偏倚与覆盖率不足,意向治疗与符合方案效应随切换程度增大而分化;一项心脏移植分析验证了这些结论。在第二种机制中,四项发现中的三项可重复,排名反转减弱,模型敏感性依赖于校准情况。
英文摘要
On a non-collapsible survival mechanism, longitudinal-matching and target-trial-emulation methods are not competing estimators of one truth but answers to different causal questions, so a benchmark that scores them against a single "true hazard ratio" fabricates bias. We provide the direct comparison of relative efficiency, variance estimation, and model sensitivity that reviews find lacking. On a deliberately non-collapsible continuous-time Cox mechanism with known truth, the dominant families (sequential Cox, sequential stratification, risk-set matching, and inverse-probability-of-treatment-weighted (IPTW) marginal structural models) target numerically distinct causal estimands (marginal, conditional, two average-treatment-effect-on-the-treated, and intention-to-treat versus per-protocol). First, we quantify the phantom bias a shared marginal truth fabricates: 0.32-0.33 log-cumulative-hazard-ratio units for the matching estimators and 0.15 for the conditional method; the associational naive time-dependent Cox sits 0.76 away, a total discrepancy compounding the estimand gap with confounding. Second, a rank reversal: the recommended method flips with the target estimand, and a low-variance off-target estimator can still win on mean-squared error. Third, a cross-family variance result: the cluster-robust sandwich is closer to nominal for the trial-stacking estimator (0.90) but under-covers the matching estimators (0.77-0.82), which a prespecified n=500 bootstrap sub-study brings to 0.95-0.96. Fourth, model sensitivity: omitting a confounder induces 0.45-0.50 log-hazard-ratio bias and undercoverage, and intention-to-treat and per-protocol effects diverge as switching increases; a heart-transplant analysis illustrates these. On a second mechanism three of four findings replicate, the rank reversal attenuating and model sensitivity proving calibration-dependent.
Comments19 pages, 2 figures, 7 tables; 28-page supplementary appendix; reproducible code: https://github.com/ehsanx/lmtte