当治疗数据缺失时学习治疗对象
Learning Who to Treat When Treatment is Missing
浏览论文内容
中文总结 AI 辅助
研究在治疗数据缺失时的策略学习问题,通过扩展有效估计器到MAR和MCCAR数据下的策略值及CATE估计,经渐近效率分析和实验验证,证明MAR估计器更优,为从业者提供稳健策略学习工具。
中文摘要 AI 辅助
策略学习方法越来越多地用于在预算约束下指导治疗分配。大多数提出的方法假设治疗数据完整,但实际应用中经常存在缺失值,这可能会使估计产生偏差并导致次优策略。我们通过将平均治疗效果(ATE)估计的有效估计器扩展到随机缺失(MAR)和完全条件随机缺失(MCCAR)治疗数据下的策略值和条件平均治疗效果(CATE)估计来解决这一差距。通过渐近效率分析,我们证明了利用部分观测单元的MAR估计器在MCCAR假设成立时既有效又比MCCAR估计器更有效。我们使用合成和半合成数据集进行的综合实验证实,正确指定缺失机制至关重要:错误指定的估计器无论样本量如何都会存在偏差,而我们的估计器在假设满足时可实现接近最优的性能。我们的工作为从业者提供了理论基础扎实、经实证验证的工具,用于在存在缺失治疗数据的情况下进行稳健的策略学习。
英文摘要
Policy learning methods are increasingly used to inform treatment allocation under budget constraints. Most proposed methods assume complete treatment data, yet applications frequently suffer from missingness that can bias estimates and lead to suboptimal policies. We address this gap by extending efficient estimators for average treatment effect (ATE) estimation to policy value and conditional average treatment effect (CATE) estimation under missing at random (MAR) and missing completely conditionally at random (MCCAR) treatment data. Through asymptotic efficiency analysis, we prove that the MAR estimator, which leverages partially-observed units, is both valid and more efficient than the MCCAR estimator when MCCAR assumptions hold. This result provides formal justification for preferring MAR-based estimation in policy learning under both missing data settings. Our comprehensive experiments using synthetic and semi-synthetic datasets confirm that correctly specifying the missingness mechanism is crucial: misspecified estimators remain biased regardless of sample size, while our estimators achieve near-oracle performance when assumptions are satisfied. Our work provides practitioners with theoretically grounded, empirically validated tools for robust policy learning in the presence of missing treatment data.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。