arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00284cs.LGcs.AIstat.ML

从正例-未标注数据中最大化部分AUC

Partial AUC Maximization from Positive-unlabeled Data

Atsutoshi Kumagai, Tomoharu Iwata, Taishi Nishiyama, Hiroshi Takahashi, Kazuki Adachi, Yasuhiro Fujiwara

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种仅从正例和未标注数据中最大化部分AUC的方法,无需负例数据,通过经验风险最小化框架推导估计器并实验验证其有效性。

中文摘要 AI 辅助

受试者工作特征曲线下的部分面积(pAUC)是二分类中一个重要的性能指标,它总结了在特定假阳性率(FPR)范围内真阳性率的概况。在网络安全、医疗保健和广告等许多实际应用中,需要获得实现高pAUC的分类器。尽管已经提出了许多最大化pAUC的方法,但它们通常需要同时标注正例和负例数据进行训练。然而,在实践中,由于隐私问题或标注需要高度专业知识,标注负例数据往往难以收集。在本文中,我们提出了一种从正例和未标注(PU)数据中最大化pAUC而无需负例数据的方法。在经验风险最小化框架内,我们证明了pAUC(包括其依赖于FPR的阈值)可以仅使用正例密度和边际密度来表示,并推导出基于PU数据的经验估计器。然后通过最大化所推导的平滑经验pAUC估计器来训练分类器。我们通过十个真实世界数据集实验证明了所提出方法的有效性。

英文摘要

The partial area under the receiver operating characteristic curve (pAUC) is an important performance metric for binary classification that summarizes true positive rates within a specific range of false positive rates (FPRs). Classifiers that achieve high pAUC need to be obtained in many real-world applications such as cybersecurity, medical care, and advertising. Although many methods for maximizing the pAUC have been proposed, they typically require both labeled positive and negative data for training. However, in practice, labeled negative data are often difficult to collect due to privacy concerns or the need for high expertise to annotate them. In this paper, we propose a method for maximizing the pAUC from positive and unlabeled (PU) data without negative data. Within an empirical risk minimization framework, we show that the pAUC, including its FPR-dependent thresholds, can be represented using only the positive and marginal densities, and derive an empirical estimator from PU data. A classifier is then trained by maximizing the derived smoothed empirical pAUC estimator. We experimentally demonstrate the effectiveness of the proposed method with ten real-world datasets.

发表机构

  • NTT, Inc.(日本电信电话公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑