arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估认知不确定性:超越分布外检测和主动学习

Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning

Jakub Paplhám, Willem Waegeman, Eyke Hüllermeier, Vojtěch Franc

arXiv 2607.14817首次发表:更新:

发表机构

Czech Technical University in Prague; Ghent University; LMU Munich; MCML; DFKI(布拉格捷克技术大学; 根特大学; 慕尼黑大学; 机器学习与数据挖掘中心; 德国人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究认知不确定性评估,基于认知拒绝选项框架,通过将选择性预测表述为约束优化证明最优选择器,揭示现有文献弱点,提议评估分解的可实现风险等作为诊断,实验表明决策理论排名与代理任务排名有差异。

AI 中文摘要

当前对认知不确定性的评估依赖于分布外检测和主动学习等任务。然而,这些任务的贝叶斯最优决策策略与常用于量化认知不确定性的分数不一致。基于认知拒绝选项框架,我们利用其识别遗憾(可减少误差)的能力来评估认知不确定性。将选择性预测表述为对覆盖率、预期风险和遗憾的约束优化,我们证明最优选择器是真实偶然和认知不确定性的阈值化凸组合。这种理论统一揭示了近期不确定性解缠文献中的一个弱点:我们表明学习组件之间的标准相关性指标不一定能预测它们的实际操作效用。我们反而提议评估分解的可实现风险、遗憾、覆盖表面,作为联合解缠和效用的诊断方法。在具有密集人工标注的数据集上对标准方法进行基准测试表明,决策理论排名可能与代理任务排名有很大差异,包括在一个标准上排名靠前而在另一个标准上排名靠后的方法之间的成对排名反转。

英文摘要

Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable risk, regret, coverage surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑