高维因果外推的风险校准平衡
Risk-Calibrated Balancing for High-Dimensional Causal Extrapolation
浏览论文内容
中文总结 AI 辅助
针对高维弱重叠下协变量平衡导致方差增大的问题,提出风险校准平衡方法,通过岭惩罚和条件预测风险选择参数,在模拟和真实数据中降低目标风险并改进估计。
中文摘要 AI 辅助
在观察性因果推断中,协变量平衡被广泛用于减少源-目标协变量偏移,但在高维弱重叠情况下,更强的平衡可能导致权重集中并增加方差。平衡衡量目标协变量分布被表示的程度,但其本身并不能确定反事实均值能被多可靠地估计。我们针对处理组平均处理效应开发了风险校准平衡方法,该方法对任何归一化的基础权重应用岭增强,并利用反事实均值的条件预测风险来选择其惩罚参数。在随机效应预测模型下,我们推导出该风险精确的有限样本分解,将其分为残差协变量不平衡和权重引起的方差两部分。对于比例渐近下的设计无关基础权重,我们刻画了极限风险如何依赖于源和目标协方差几何、总体均值偏移以及权重集中度。对于协变量自适应基础权重,我们开发了一个一致的目标感知风险估计器,其最小化器达到消失的缩放预言机超额风险。模拟表明,高维风险预测对自适应平衡仍然具有信息量,且目标感知调参通常能降低超额目标风险。对职业培训和单细胞扰动数据的实证分析显示,风险校准平衡通常在相应基础估计器上有所改进,且在弱重叠情况下改进更大。
英文摘要
In observational causal inference, covariate balancing is widely used to reduce source-target covariate shift, but under weak overlap in high dimensions, stronger balance can induce concentrated weights and increase variance. Balance measures how well the target covariate distribution is represented, but does not by itself determine how reliably the counterfactual mean can be estimated. We develop risk-calibrated balancing for the average treatment effect on the treated, which applies ridge augmentation to any normalised base weights and selects its penalty using conditional prediction risk of the counterfactual mean. Under a random-effects predictive model, we derive an exact finite-sample decomposition of this risk into residual covariate imbalance and weight-induced variance. For design-independent base weights under proportional asymptotics, we characterise how limiting risk depends on source and target covariance geometry, population mean shift, and weight concentration. For covariate-adaptive base weights, we develop a uniformly consistent target-aware risk estimator whose minimiser attains vanishing scaled oracle excess risk. Simulations show that the high-dimensional risk predictions remain informative for adaptive balancing and that target-aware tuning generally reduces excess target risk. Empirical analyses of job-training and single-cell perturbation data show that risk-calibrated balancing generally improves on the corresponding base estimators, with larger gains under weaker overlap.