AI 中文总结
该研究针对分布偏移问题,提出AIHW方法,通过分离系统偏移与残余随机扰动并权衡两类不确定性,在多站点数据集上验证其可降低均方误差、提升覆盖范围,具有良好鲁棒性。
AI 中文摘要
对源样本进行重加权以匹配目标协变量分布,是将证据从一个总体推广到另一个总体时应对分布偏移的标准策略。该策略非常适合确定性、可学习的协变量差异,但当源-目标总体差异还包含协变量偏移之外的变化,或密度比权重的估计不稳定时,可能不够用。为应对这一挑战,我们引入了一种新模型,该模型在考虑系统偏移后,允许两个总体分布之间存在非系统性变化。这种残余偏移被建模为概率空间的随机扰动,无法以可学习的方式表示。通过这种方式,我们将系统偏移(视为偏差,通过重加权校正)与残余随机扰动(视为分布不确定性,通过数据集池化处理)区分开来。在纯随机扰动下,该原理产生了增强型逆距离加权(AIDW),其使用回归增强和方差最优的数据集级池化。对于混合偏移,我们开发了增强型逆混合加权(AIHW),它在AIDW和标准增强型重要性加权之间进行插值。两种方法都通过一个描述随机扰动强度的“分布距离”来权衡抽样不确定性和分布不确定性。我们建立了这些方法的渐近性质,以及选择调优参数的插件指导和模型诊断工具。在三个真实世界多站点数据集上的实验表明,与标准加权基线相比,均方误差持续降低,且在仅协变量偏移调整会导致覆盖不足的场景中,经验覆盖范围显著改善,显示了所提方法在不同分布偏移场景下的鲁棒性。
英文摘要
Reweighting source samples to match a target covariate distribution is a standard response to distribution shift when generalizing evidence from one population to another. This strategy is well suited to deterministic, learnable covariate discrepancies, but can be insufficient when source--target population differences also contain changes beyond covariate shift or when estimation of the density-ratio weights is unstable. To address this challenge, we introduce a new model that allows non-systematic changes between two population laws after systematic shifts are accounted for. Such residual shift is modeled as random perturbations to the probability space that cannot be represented in a learnable way. In this way, we separate systematic shifts, treated as bias and corrected by reweighting, from residual random perturbations, treated as distributional uncertainty and handled through dataset pooling. Under pure random perturbations, this principle yields Augmented Inverse Distance Weighting (AIDW), which uses regression augmentation and variance-optimal dataset-level pooling. For mixed shifts, we develop Augmented Inverse Hybrid Weighting (AIHW), which interpolates between AIDW and standard augmented importance weighting. Both methods trade off sampling uncertainty and distributional uncertainty via a \emph{distributional distance} that describes the strength of random perturbations. We establish asymptotic properties of the methods, together with plug-in guidance for choosing tuning parameters and model diagnostic tools. Experiments on three real-world multi-site datasets demonstrate consistent reductions in mean-squared error compared with standard weighting baselines, along with substantially improved empirical coverage in settings where covariate-shift adjustment alone undercovers, showing the robustness of the proposed methods across diverse distribution shift scenarios.