通过模型平均处理协变量偏移
Handling covariate shift by model averaging
浏览论文内容
中文总结 AI 辅助
针对协变量偏移问题,本文提出自适应重要加权模型平均方法,通过构造不同指数的重要加权最小二乘估计量并取数据驱动平均,在模拟与真实数据应用中实现了良好的目标预测性能。
中文摘要 AI 辅助
在现代数据分析中,构建统计方法所用数据与最终应用总体之间的分布不匹配现象十分普遍。我们研究该问题的一个基本实例——协变量偏移,并开发一种自适应重要加权模型平均方法,用于当源分布中有标记观测值、目标分布中仅观测到未标记协变量时的预测任务。直接基于源样本拟合的方法通常在源分布下优化预测风险,因此可能对目标预测并非最优。通过目标与源协变量边际的密度比进行重要加权是一种自然的校正方式,但少量较大的密度比值会在有限样本中大幅增大所得估计量的方差。我们将重要加权校正的程度视为模型不确定性的来源,以解决这一偏差-方差权衡问题。具体而言,我们通过将估计的密度比提升至一系列指数来构造自适应重要加权最小二乘估计量族,其端点分别对应普通最小二乘和标准重要加权最小二乘,并对这些候选估计量形成数据驱动的平均。在模型误设情况下,所提模型平均估计量相对于候选估计量不可行的最优凸组合,被证明是渐近最优的;在模型正确设定情况下,发散的惩罚项会使选定的权重集中在普通最小二乘端点附近。模拟和真实数据应用表明,所提方法在考虑的各类设置下均能达到具有竞争力的目标预测性能。
英文摘要
Distributional mismatch between the data used to construct a statistical procedure and the population to which it is ultimately applied is pervasive in modern data analysis. We study covariate shift, a fundamental instance of this problem, and develop an adaptive importance-weighted model averaging method for prediction when labeled observations are available from a source distribution, whereas only unlabeled covariates are observed from the target distribution. Procedures fitted directly to the source sample generally optimize prediction risk under the source distribution and may therefore be suboptimal for target prediction. Importance weighting by the density ratio between the target and source covariate marginals provides a natural correction, but a small number of large density-ratio values can substantially inflate the variance of the resulting estimator in finite samples. We address this bias-variance trade-off by treating the degree of importance-weighting correction as a source of model uncertainty. Specifically, we construct a family of adaptive importance-weighted least-squares estimators by raising the estimated density ratio to a range of exponents, with the endpoints corresponding to ordinary least squares and standard importance-weighted least squares, and form a data-driven average over these candidates. Under model misspecification, the proposed model averaging estimator is shown to be asymptotically optimal relative to the infeasible best convex combination of the candidate estimators. Under correct specification, a diverging penalty is shown to make the selected weights concentrate near the ordinary least-squares endpoint. Simulations and a real-data application show that the proposed method achieves competitive target-prediction performance across the settings considered.