arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReliableNet:深度学习中可信分类的机会约束方法

ReliableNet: A Chance-Constrained Approach to Trustworthy Classification in Deep Learning

Ange-Clément Akazan, Ineza Remy Mugenga, Abebe Geletu, Jean Medard Ngnotchouye, Issa Karambal

arXiv 2608.09768首次发表:更新:

发表机构

School of Agriculture and Science, University of KwaZulu-Natal; AIMS Research and Innovation Centre, African Institute for Mathematical Sciences; African Institute for Mathematical Sciences(夸祖鲁-纳塔尔大学农业与科学学院; 非洲数学科学研究所AIMS研究与创新中心; 非洲数学科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReliableNet是一种将联合自信-错误概率约束在用户指定风险预算内的深度学习可信分类方法,在各类数据偏移下均表现优异。

AI 中文摘要

既自信又错误的预测是关键的可靠性故障,因为当模型出错时,它会绕过弃权(不执行)和人工审核。经验风险最小化(ERM)控制平均损失,但不直接控制这种故障;而校准、不确定性估计、共形风险控制和选择性预测方法针对的是相关的可靠性属性,而非在训练期间对联合故障事件进行约束。我们提出ReliableNet,它将联合自信-错误(JCW)概率(即预测同时具有自信和错误的概率)约束在用户指定的风险预算α∈(0,1)以下。我们将其表述为一个机会约束的ERM问题,使用保守的平滑内近似,其总体可行性意味着原始的JCW约束。在四个表格数据集和两个图像数据集上,与ERM、事后校准、共形风险控制和选择性预测等基线方法相比,ReliableNet是唯一在分布内针对每个数据集和随机种子都能在JCW预算内得到验证的方法。在人口统计、歧义、虚假相关性、新类别和协变量偏移下,它在保持与准确率、覆盖率、校准和选择性预测极具竞争力的同时,实现了对比方法中最低的经验JCW。风险-覆盖率结果进一步表明,ReliableNet在大多数数据集上比基准方法实现了更好的选择性排序。总体而言,ReliableNet为可信分类提供了一种原则性方法。

英文摘要

A prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken. Empirical risk minimization (ERM) controls average loss but not this failure directly, while calibration, uncertainty estimation, conformal risk control, and selective prediction methods target related reliability properties rather than bounding the joint failure event during training. We propose ReliableNet, which constrains the Joint Confident-Wrong (JCW) probability, the probability that a prediction is simultaneously confident and incorrect, below a user-specified risk budget $α\in(0,1)$. We formulate this as a chance-constrained ERM problem, use a conservative smooth inner approximation whose population feasibility implies the original JCW constraint. Across four tabular and two image datasets, ReliableNet is the only method certified within the JCW budget for every dataset and seed in distribution, when compared against baselines spanning ERM, post-hoc calibration, conformal risk control, and selective prediction. Under demographic, ambiguity, spurious-correlation, novel-class, and covariate shifts, it achieves the lowest empirical JCW among the compared methods while remaining very competitive in accuracy, coverage, calibration, and selective prediction. Risk-coverage results further indicate that ReliableNet achieves better selective ranking than the benchmark methods on most datasets. Overall, ReliableNet provides a principled approach to trustworthy classification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑