发表机构
Yale School of Public Health(耶鲁公共卫生学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对因果推断中预测误差作为干扰函数估计度量的不足,通过模拟对比多种方法,发现预测误差与因果性能关联不稳定,联合误差度量也不适用,建议谨慎使用预测误差评估因果估计质量。
AI 中文摘要
预测误差被广泛用于评估因果推断中的干扰函数估计量,但其与因果估计量性能的关系可能因性能度量的不同而存在差异。我们在部分线性模型中通过蒙特卡洛模拟研究了该问题,比较了普通最小二乘法(OLS)、广义加性模型(GAMs)、XGBoost以及结合XGBoost的双重机器学习(DML-XGBoost),评估了干扰函数的预测误差、偏差、均方根误差(RMSE)以及95%置信区间覆盖率。我们还考察了一种基于暴露和结果干扰函数估计误差绝对交叉乘积的简单联合误差度量。在所有模拟场景中,XGBoost在非理想方法中具有最低的RMSE,而DML-XGBoost通常提供更好的置信区间覆盖率。预测误差并未在所有方法和场景中一致跟踪因果偏差,且点估计性能最佳的方法未必具有最佳的置信区间覆盖率。联合误差度量与因果偏差仅存在弱关联,无法作为因果性能的独立有效度量。这些结果表明,预测误差可用于评估干扰函数估计,但不应被视为所得因果估计量质量的直接度量。
英文摘要
Prediction error is widely used to evaluate nuisance-function estimators in causal inference, but its relationship with causal estimator performance may differ across performance measures. We studied this question in a partially linear model using Monte Carlo simulations. We compared ordinary least squares (OLS), generalized additive models (GAMs), XGBoost, and Double Machine Learning with XGBoost (DML-XGBoost), evaluating nuisance-function prediction error, bias, RMSE, and 95\% confidence interval coverage. We also examined a simple joint-error measure based on the absolute cross-product of estimation errors from the exposure and outcome nuisance functions. Across the simulated settings, XGBoost had the lowest RMSE among the non-oracle methods, while DML-XGBoost generally provided better confidence interval coverage. Prediction error did not consistently track causal bias across methods and settings, and the method with the best point-estimation performance did not necessarily have the best confidence interval coverage. The joint-error measure was only weakly associated with causal bias and did not provide a useful standalone measure of causal performance. These results suggest that prediction error is useful for assessing nuisance-function estimation, but it should not be treated as a direct measure of the quality of the resulting causal estimator.
Comments10 pages, 2 figures, 1 table