双重机器学习的解析型与自举置信区间:模拟研究及在城乡肥胖患病率差异中的应用
Analytical and Bootstrap Confidence Intervals for Double Machine Learning: Simulation Studies and an Application to Rural-Urban Differences in Obesity Prevalence
浏览论文内容
中文总结 AI 辅助
该研究通过模拟和真实数据分析,对比了双重机器学习的解析型与自举置信区间性能,发现学习器选择影响推断可靠性,样本量增加时两类置信区间覆盖概率下降,且乡村化程度显著提升美国县域肥胖患病率。
中文摘要 AI 辅助
双重机器学习(Double Machine Learning, DML)是适用于多种场景的流行处理效应估计方法,它允许使用多种灵活的机器学习方法估计干扰参数,同时保证推断的有效性。但在实践中,应用研究者需从众多机器学习算法中选择干扰模型,而该选择对DML方差估计的影响尚未得到充分表征。我们开展了全面的模拟研究,比较不同机器学习算法下DML置信区间的覆盖概率,研究中对比了两类置信区间:(1)基于DML理论推导的解析型置信区间;(2)自举置信区间。我们在不同数据生成设置下使用的学习器包括普通最小二乘法、LASSO、随机森林、LightGBM和神经网络,通过偏差、置信区间宽度,以及最关键的覆盖概率来评估不同设置下的性能。结果显示,解析型与自举置信区间的覆盖性能存在显著差异,凸显学习器选择对可靠DML推断至关重要;令人惊讶的是,我们发现许多场景中,当样本量增加时,DML解析型和自举置信区间的覆盖概率均会下降。我们进一步使用美国各县城乡差异的真实数据集研究覆盖概率,真实数据分析发现:(1)模型性能仍因学习器选择而异;(2)更高的乡村化程度对县级肥胖患病率具有统计学显著的正向影响。
英文摘要
Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine learning methods to be used for nuisance parameter estimation while preserving valid inference. In practice, however, applied researchers must choose among many machine learning algorithms for nuisance models, and the impact of this choice on the variance estimation of DML is not well characterized. We conduct a comprehensive simulation study to compare the coverage probability of DML confidence intervals across different machine learning algorithms. In this study, we compare (1) analytical confidence intervals derived by DML theory versus (2) bootstrap confidence interval. We use a set of learners including ordinary least squares, LASSO, Random Forest, LightGBM, and Neural Networks under different data generation settings. We evaluate the performance across difference settings by bias, confidence interval width, and most importantly, coverage probability. Our results show substantial variability in coverage performance across analytical and bootstrap confidence intervals, highlighting that learner choice plays a critical role in reliable DML inference. Surprisingly, we find that in many settings, when sample size increases, the coverage probability of both DML analytical and bootstrap confidence interval decreases. We further investigate coverage probabilities using a real dataset on rural urban differences among U.S. counties. The real data analysis discovers that (1) the model performance still varies by the learner choices and (2) greater rurality has a statistically significant increasing effect on county level obesity prevalence.
发表机构
- Vanderbilt University Medical Center(范德堡大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。