协变量偏移下代理衍生模型的推断评估
Inferential Evaluation of Surrogate-Derived Models under Covariate Shift
查看机构详情
- National University of Singapore(新加坡国立大学)
- Peking University Health Science Center, Peking University(北京大学医学部,北京大学)
- Beijing International Center for Mathematical Research, Peking University(北京大学北京国际数学研究中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对迁移学习中协变量偏移下金标签稀缺的问题,提出交叉拟合估计量等方法,实现代理衍生模型在目标人群的TPR、FPR等指标的推断,并通过模拟与实际AI应用验证了方法有效性。
中文摘要 AI 辅助
在迁移学习场景中,从充足代理标签衍生的模型可能被部署到未观测到金标准结果的目标人群中。评估其目标性能对确定基于该模型的决策是否仍可靠至关重要,但当金标签稀缺且不同数据源的协变量分布存在差异时,这一评估颇具挑战。我们研究了三样本场景,包含少量带金标签的源数据、大量带代理标签的源数据以及无标签的目标数据。在条件可迁移性假设下,我们针对目标人群中潜在的金标准结果评估代理衍生模型。我们提出了交叉拟合估计量,通过特定于源数据的密度比率从两个带标签源数据中迁移信息;还结合结果回归增强与核校正,以估计阈值附近的模型,同时考虑三个样本带来的不确定性。我们建立了真阳性率(TPR)和假阳性率(FPR)的渐近线性推断、受试者工作特征(ROC)曲线的一致性及逐点推断,以及曲线下面积(AUC)的渐近正态推断。模拟实验评估了偏差、覆盖率以及对带宽和相对样本量的敏感性;对Chatbot Arena的回顾性时间验证和半合成ACS-Income研究则在现实世界AI应用中提供了验证。
英文摘要
In transfer-learning settings, a model derived from abundant surrogate labels may be deployed in a target population where gold-standard outcomes are unobserved. Evaluating its target performance is essential for determining whether decisions based on the model remain reliable, yet it is difficult when gold labels are scarce, and covariate distributions differ across data sources. We study a three-sample setting with a small gold-labeled source, a larger surrogate-labeled source, and an unlabeled target. Under conditional transportability, we evaluate the surrogate-derived model against the latent gold-standard outcome in the target population. We propose cross-fitted estimators that transport information from the two labeled sources through source-specific density ratios. We also combine outcome-regression augmentation with a kernel correction for estimating the model near a threshold, accounting for uncertainty from all three samples. We establish asymptotically linear inference for TPR and FPR, consistency and pointwise inference for the ROC curve, and asymptotically normal inference for AUC. Simulations assess bias, coverage, and sensitivity to bandwidth and relative sample sizes. A retrospective temporal validation on Chatbot Arena and a semi-synthetic ACS-Income study provide validation in real-world AI applications.