发表机构
Google, Inc.; Department of Biostatistics, University of Michigan; Department of Statistics, University of California Santa Barbara(谷歌公司; 密歇根大学生物统计学系; 加州大学圣塔芭芭拉分校统计系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对部署时辅助信息不可用的问题,提出ALCATRAs框架,通过成本自适应任务选择与替代学习,在预算约束下高效获取辅助信息以改进预测,并证明其理论有效性与样本效率优势。
AI 中文摘要
许多科学研究允许在数据标注过程中收集昂贵的辅助信息,但在部署时却无法获取这些信息。例如诊断测试、实验室化验和专家评估。我们研究这种部署不对称性下的预测问题,其中辅助变量在预算约束下于标注阶段被选择性获取,但在预测时系统性地不可用,从而产生了一个由设计导致的缺失问题,该问题将数据获取、替代模型构建和预测耦合在一起。在这项工作中,我们引入了具有成本自适应任务资源分配(ALCATRAs)的主动学习,这是一个在资源约束下选择性获取辅助信息并利用这些信息改进下游预测的统一框架。ALCATRAs由两个主要组成部分构成:一个任务选择策略,它策略性地为未标注数据选择一系列成本效益高的任务来执行;以及一个替代学习过程,它从已完成的任务中迁移知识以增强模型预测。在理论上,我们证明了替代模型和样本高效任务策略在改进模型预测误差界方面的有效性。模拟研究和在UCI心脏病队列上的应用表明,在所研究的设置下,所提出的ALCATRAs框架相对于基线方法具有更高的样本效率。
英文摘要
Many scientific studies allow costly auxiliary information to be collected during data labeling but not at deployment. Examples include diagnostic tests, laboratory assays, and expert evaluations. We study prediction under this deployment asymmetry, where auxiliary variables are selectively acquired during labeling under a budget constraint but systematically unavailable at prediction time, creating a missing-by-design problem that couples data acquisition, surrogate construction, and prediction. In this work, we introduce Active Learning with Cost-Adaptive Task Resource Allocations (ALCATRAs), a unified framework for selectively acquiring auxiliary information under resource constraints and leveraging that information to improve downstream prediction. ALCATRAs consists of two main components: a task-selection policy which strategically selects a sequence of cost-effective tasks for unlabeled data to perform, and a surrogate learning procedure which transfers knowledge from completed tasks to enhance model predictions. In theory, we show the effectiveness of surrogate models and sample-efficient task policies in improving the model's prediction error bound. Simulation studies and an application to the UCI heart disease cohort demonstrate improved sample efficiency of the proposed ALCATRAs framework relative to baselines under the studied settings.