AI 中文总结
针对部署人群与训练人群不同导致的预测性能下降,提出SAGE一步估计器,利用目标子组摘要更新源训练预测器,降低平均渐近目标超额风险,并在真实数据上优于忽略偏移的更新和EB方法。
AI 中文摘要
当部署人群与训练人群不同时,预测模型可能表现不佳。目标人群的数据会有所帮助,但由于访问限制或报告惯例,个体层面的目标数据可能无法获取。我们考虑一个多分辨率设置,其中个体层面数据可从源人群获得,而目标人群仅通过子组摘要观测。我们提出SAGE,一种一步估计器,利用从这些摘要估计的梯度更新源训练预测器。受与随机分布偏移模型一致的诊断的启发,我们选择SAGE的步长以同时考虑抽样和分布不确定性。在该模型下,SAGE相对于源训练预测器降低了平均渐近目标超额风险。我们还表明,在该模型下,熵平衡加权经验风险最小化(EB),即重新加权源观测以匹配目标摘要,渐近等价于全步SAGE更新。具有最优步长的SAGE的渐近均方误差不大于EB。在真实世界数据集上,SAGE通常优于源训练预测器和忽略分布偏移的一步更新,包括当随机偏移模型仅部分捕捉观测到的偏移时。与EB相比,SAGE在研究的样本量和偏移设置中更一致地改善预测。
英文摘要
Prediction models can perform poorly when the deployment population differs from the training population. Data from the target population would help, but individual-level target data may be inaccessible because of access restrictions or reporting conventions. We consider a multi-resolution setting in which individual-level data are available from a source population, while the target population is observed only through subgroup summaries. We propose SAGE, a one-step estimator that updates a source-trained predictor using a gradient estimated from these summaries. Motivated by diagnostics consistent with the random distribution shift model, we choose SAGE's step-size to account for both sampling and distributional uncertainty. Under this model, SAGE reduces mean asymptotic target excess risk relative to the source-trained predictor. We also show that, under the model, entropy-balancing weighted empirical risk minimization (EB), which reweights source observations to match the target summaries, is asymptotically equivalent to a full-step SAGE update. SAGE with the optimal step-size has asymptotic mean squared error no larger than that of EB. Across real-world datasets, SAGE generally improves on the source-trained predictor and one-step updates that ignore distribution shift, including when the random shift model only partially captures the observed shifts. Compared to EB, SAGE improves prediction more consistently across the sample size and shift settings studied.