arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

贝叶斯最优误码率与受试者工作特征曲线下面积:估计与估计器评估

Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators

Ryota Ushio, Takashi Ishida, Masashi Sugiyama

arXiv 2609.02304首次发表:更新:

发表机构

The University of Tokyo; RIKEN AIP(东京大学; 理化学研究所人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对类别不平衡或标注噪声场景,提出基于软标签的最优BER和AUC估计器,并扩展FeeBee框架用于评估,经实验验证其有效性。

AI 中文摘要

机器学习中的一个基本量是任意模型在给定任务上可达到的最优性能。估计该量有助于我们将误差的不可约部分与模型的不足区分开来,从而明确剩余的改进空间。近期研究表明,在二分类任务中可通过软标签估计贝叶斯误差(即等价的最优准确率)。然而,在类别严重不平衡或标注存在噪声的场景中,准确率往往无法有效概括性能,此时平衡误码率(BER)和受试者工作特征曲线下面积(AUC)等指标更为适用。我们针对这一空白提出两项互补贡献:(一)估计方法:我们提出基于软标签的最优BER和AUC估计器。首先考虑真实软标签和类别先验已知的干净场景,随后将估计器扩展至更现实的场景——类别先验未知,且观测到的软标签受未知保序变换(可能后续叠加加性噪声)污染。在后一场景中,我们通过结合辅助硬标签的保序回归近似恢复干净软标签,通过硬标签的截断均值估计类别先验,并推导所得插件估计器的有限样本误差界。(二)评估方法:由于最优值在真实数据集上不可观测,评估此类估计器本身颇具挑战性。我们将原本为评估贝叶斯误差估计器提出的FeeBee框架扩展至最优BER和AUC,所得流程无需知晓最优值即可提供实用评估分数,且适用于任意最优BER或AUC估计器,不仅限于我们提出的估计器。在合成数据集和真实世界数据集上的实验验证了所提估计器和评估流程的有效性。

英文摘要

A fundamental quantity in machine learning is the optimal performance achievable by any model on a given task. Estimating this quantity allows us to distinguish the irreducible part of the error from a deficiency of the model, telling us how much room for improvement remains. Recent work has shown that the Bayes error, or equivalently the optimal accuracy, can be estimated from soft labels in binary classification. However, accuracy is often a poor summary of performance in settings with severe class imbalance or noisy annotations, where metrics such as the balanced error rate (BER) and the area under the ROC curve (AUC) are more appropriate. We address this gap with two complementary contributions. (i) Estimation. We propose soft-label-based estimators for the optimal BER and AUC. We first consider the clean setting in which true soft labels and the class prior are known, and then extend the estimators to a more realistic setting in which the class prior is unknown and the observed soft labels are corrupted by an unknown order-preserving transformation, possibly followed by additive noise. In the latter setting, we approximately recover the clean soft labels via isotonic regression with auxiliary hard labels, estimate the class prior with a clipped mean of the hard labels, and derive finite-sample error bounds for the resulting plug-in estimators. (ii) Evaluation. Since the optimum is unobservable on real datasets, evaluating any such estimator is itself nontrivial. We extend the FeeBee framework, originally proposed for evaluating Bayes-error estimators, to the optimal BER and AUC. The resulting procedure provides practical evaluation scores without requiring knowledge of the optimum, and applies to any estimator of the optimal BER or AUC, not only our proposed ones. Experiments on synthetic and real-world datasets validate both the estimators and the evaluation procedure.

Comments43 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑