基于置信度的有偏正类-未标注数据AUC最大化
AUC Maximization from Biased Positive-unlabeled Data with Confidence
浏览论文内容
中文总结 AI 辅助
针对有偏正类-未标注数据,提出利用置信度估计AUC风险,实现AUC最大化,并在八个真实数据集上验证有效性。
中文摘要 AI 辅助
最大化接收者操作特征曲线下面积(AUC)是不平衡二分类问题的标准方法。尽管最大化AUC需要正类和负类数据,但在某些实际应用中,由于隐私问题或需要专业知识进行标注,负类数据往往难以收集。因此,从正类和未标注(PU)数据中最大化AUC已引起关注。现有方法假设标注的正类数据是真实正类分布的无偏样本。然而,这一理想假设在实践中经常被违反。本文提出了一种从有偏PU数据中最大化AUC的方法。为解决偏差问题,我们的关键思想是利用与少量标注正类数据相关联的置信度,即实例为正类的概率。我们利用带有置信度的有偏PU数据推导出AUC风险的估计器,从而在此类偏差下实现AUC最大化。我们进一步证明,当可用置信度是真实后验概率的任意严格递增变换时,重写的AUC风险可诱导出贝叶斯最优的AUC排序。我们在八个真实世界数据集上通过实验证明了我们方法的有效性。
英文摘要
Maximizing the area under the receiver operating characteristic curve (AUC) is a standard approach to imbalanced binary classification. Although positive and negative data are required for maximizing the AUC, negative data are often difficult to collect in some real-world applications due to privacy concerns or the need for specialized expertise to annotate them. Thus, AUC maximization from positive and unlabeled (PU) data has been attracting attention. Existing methods assume that labeled positive data are unbiased samples from the true positive distribution. However, this ideal assumption is often violated in practice. In this paper, we propose a method to maximize the AUC from biased PU data. To address the bias, our key idea is to exploit {\it confidence}, i.e., the probability that an instance is positive, associated with the small number of labeled positive data. We derive an estimator of the AUC risk using biased PU data with confidence, enabling AUC maximization under such bias. We further show that the rewritten AUC risk induces a Bayes-optimal AUC ranking even when the available confidence is any strictly increasing transformation of the true posterior probability. We experimentally show the effectiveness of our method on eight real-world datasets.
发表机构
- NTT, Inc.(日本电信电话公司)
机构由 AI 辅助整理,请以论文原文为准。