发表机构
Tampere university(坦佩雷大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出极端二分类问题,要求假阴性率随正例数趋于零,基于极值理论设计阈值自适应与置换检验特征选择,在四个真实数据集上优于现有方法,并展示了可解释性。
AI 中文摘要
虽然二分类是机器学习中研究最广泛的问题之一,但以几乎为零的假阴性率为目标来学习分类器的场景在很大程度上仍未得到探索。在本文中,我们引入了极端二分类问题,其目标是学习一个分类器,使其假阴性率 $\alpha$ 受到 $\epsilon_{N_1}=o_{N_1\to\infty}(1/N_1)$ 的约束,其中 $N_1$ 表示训练集中正例的数量。为了解决这个问题,我们提出了一种阈值自适应方法,该方法在理论上基于极值理论推导出的保证,并辅以基于样本最大值上的置换检验的特征选择过程。在四个不同规模的真实世界数据集上的实验结果表明,我们的方法与最先进的方法相比具有优势。此外,我们通过对一个癌症筛查数据集的应用展示了其可解释性。
英文摘要
While binary classification is one of the most extensively studied problems in machine learning, the regime in which the goal is to learn a classifier with an almost zero false negative rate remains largely unexplored. In this paper, we introduce the Extreme Binary Classification problem, where the objective is to learn a classifier whose false negative rate $α$ is constrained by $ε_{N_1}=o_{N_1\to\infty}(1/N_1)$, with $N_1$ denoting the number of positive examples in the training set. To address this problem, we propose a threshold adaptation method theoretically grounded in guarantees derived from Extreme Value Theory, together with a feature selection procedure based on a permutation test applied to sample maxima. Experimental results on four real-world datasets of varying sizes demonstrate that our approach compares favorably with state-of-the-art methods. In addition, we illustrate its interpretability through an application to a cancer screening dataset.