AI 中文总结
针对强处理效应异质性下标准CATE估计器失效问题,提出ICP框架,结合因果森林等技术构造有效共形区间,在高异质性合成数据和IHDP基准上验证了局部策略的精度与覆盖率优势。
AI 中文摘要
标准的条件平均处理效应(CATE)估计器在处理效应存在强异质性时会失效:条件均值的置信区间不一定覆盖个体反事实效应。我们提出了一种个性化因果预测(ICP)框架,可为特定查询单元的个体因果效应构造有限样本有效的共形预测区间。该方法利用因果森林(Causal Forest)变量重要性加权的余弦相似度将校准定位到因果相关邻域,通过合成数据扩充小局部样本,并使用满足奈曼正交性的双重稳健增强逆概率加权(AIPW)一致性得分校准区间。在标准识别假设(SUTVA和强可忽略性)以及与结果无关的校准集选择条件下,所得区间达到名义水平的边际覆盖率。该局部设计还通过使一致性得分更能代表查询单元,支持近似条件覆盖率。在高异质性合成数据集和IHDP基准上的实验表明,局部策略相较于全局基线提升了点估计精度,同时保持了名义或高于名义的覆盖率。
英文摘要
Standard CATE estimators become inadequate under strong treatment-effect heterogeneity: confidence intervals for conditional means need not cover individual counterfactual effects. We propose an Individualized Causal Prediction (ICP) framework that constructs finite-sample valid conformal prediction intervals for the individual causal effect of a specific query unit. The method localizes calibration to a causally relevant neighborhood using cosine similarity weighted by Causal Forest variable importance, augments small local samples synthetically, and calibrates intervals with doubly robust AIPW conformity scores satisfying Neyman orthogonality. Under standard identifying assumptions (SUTVA and strong ignorability) and an outcome-independent calibration-set selection condition, the resulting intervals attain marginal coverage at the nominal level. The local design also supports approximately conditional coverage by making calibration scores more representative of the query unit. Experiments on a high-heterogeneity synthetic dataset and the IHDP benchmark demonstrate that local strategies improve point accuracy over global baselines while maintaining nominal or above-nominal coverage.