发表机构
Southern University of Science and Technology; The Chinese University of Hong Kong, Shenzhen(南方科技大学; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在含噪声响应数据下的选择问题,提出鲁棒共形选择框架RCS,通过将标签噪声转化为协变量偏移问题,实现有效FDR控制,经实验验证了该方法在模拟和真实数据集上的有效性。
AI 中文摘要
共形选择已广泛应用于从大型数据集中选择高质量候选对象,并进行严格的不确定性量化,如可靠标记、药物发现和大语言模型对齐。然而,现有方法假设校准数据的响应是干净的,这在实际中很少成立。本文将上述任务表述为选择具有真实预测标签或响应超过特定值的候选对象。结果表明,在受污染的校准数据下,现有共形选择方法无法控制错误发现率(FDR)或存在严重的功效损失。为此,我们提出了鲁棒共形选择(RCS),这是一个在一般标签污染下具有有效FDR控制的选择性分类统一框架。RCS的关键在于一种新颖的统计约简:通过分别对不同类别进行条件设定,将难处理的标签噪声转化为局部协变量偏移问题,进而实现对错误选择数量的协变量调整经验贝叶斯型估计。建立了RCS的渐近FDR控制、功效最优性和鲁棒性等统计性质。我们还在随机响应模型下开发了RCS的一个实例,并将其应用于选择具有大响应值的候选对象任务。在模拟和真实世界数据集上的大量实验证明了RCS的有效性。
英文摘要
Conformalized selection has been widely applied to select high-quality candidates from large datasets with rigorous uncertainty quantification, such as reliable labeling, drug discovery, and the alignment of large language models. Nevertheless, existing methods assume clean responses on calibration data, an assumption that rarely holds in practice. In this paper, we formulate the above tasks as selecting candidates with true predicted labels or with responses exceeding certain values. We demonstrate that existing conformal selection methods fail to control the false discovery rate (FDR) or suffer from severe power loss under contaminated calibration data. To that end, we propose Robust Conformalized Selection (RCS), a unified framework for selective classification with valid FDR control under general label contamination. The key insight of RCS lies in a novel statistical reduction: by separately conditioning on different classes, we translate the intractable label noise into a localized covariate shift problem, which then enables a covariate-adjusted empirical-Bayes-type estimate of the number of false selections. Statistical properties such as the asymptotic FDR control, power optimality, and robustness of RCS are established. We further develop an instantiation of RCS under randomized response model, and also apply RCS to the task of selecting candidates with large response values. Extensive experiments on both simulated and real-world datasets demonstrate the effectiveness of RCS.