发表机构
University of Sydney Business School; University of Alberta, Department of Mathematical and Statistical Sciences(悉尼大学商学院; 阿尔伯塔大学数学与统计科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对类别重叠时两类错误不可兼得的问题,提出带弃权(不执行)的 Neyman-Pearson 分类法,通过无模型校准和鞅方法同步控制选择性两类错误,并在模拟及累犯、信用违约数据上验证。
AI 中文摘要
二分类问题通常需要在预设水平下同时控制第一类和第二类错误率。当类别分布大幅重叠时,如果每个观测都必须被分类,这两个错误限制可能不兼容。我们研究了带有弃权(不执行)的 Neyman-Pearson 分类(NPI),该方法允许不确定的观测保持未分类状态,并在每个类别内的确定性分类中评估选择性第一类和第二类错误。对于独立训练的评分,我们开发了一种无模型校准程序,该程序将所有通过校准的候选阈值对组合成一个规则,其决策区域包含每个通过候选的决策区域。我们将整个阈值搜索与每个类别内的共同鞅耦合,并精确计算其有限时域穿越概率,避免了对候选对数量的多重性调整。所得的分类器在有限样本中以用户指定的高概率同时满足两个总体选择性错误限制,且不依赖于拟合评分的准确性。我们限制了由校准引起的额外弃权(不执行),并在模拟以及累犯和信用违约预测的应用中评估了该程序。
英文摘要
Binary classification often requires controlling both type I and type II error rates at prespecified levels. When the class distributions overlap substantially, the two error limits may be incompatible if every observation must be classified. We study Neyman-Pearson classification with indecisions (NPI), which permits uncertain observations to remain unclassified and assesses selective type I and type II errors among definitive classifications within each class. For an independently trained score, we develop a model-free calibration procedure that combines all candidate threshold pairs passing calibration into a rule whose decision region contains that of every passing candidate. We couple the entire threshold search to a common martingale within each class and compute its finite-horizon crossing probability exactly, avoiding a multiplicity adjustment for the number of candidate pairs. The resulting classifier satisfies both population selective error limits simultaneously with user-specified high probability in finite samples, without assumptions on the accuracy of the fitted score. We bound the additional indecision due to calibration and evaluate the procedure in simulations and applications to recidivism and credit-default prediction.