面向可信深度学习的不确定性量化:方法与度量
Uncertainty quantification for trustworthy deep learning: Methods and measures
浏览论文内容
中文总结 AI 辅助
本综述针对深度学习的不确定性量化,梳理五大类方法及相关任务,整合评估框架,还探讨大语言模型不确定性与开放研究方向,为可信深度学习提供结构化批判性参考。
中文摘要 AI 辅助
深度神经网络在安全关键领域的部署需要可靠的预测置信度估计,但传统架构缺乏原则性的不确定性量化(UQ)。本综述对深度学习中的UQ方法进行结构化、批判性综述,范围限定为基于集成和近似贝叶斯的方法,以及用于汇总其输出的度量。与现有UQ综述相比,本文的贡献在于深入探讨高效集成近似与单遍方法,以及将生成预测分布的方法与汇总其不确定性的度量相分离的统一处理方式。我们将方法分为五大类:贝叶斯神经网络、蒙特卡洛丢弃(Monte Carlo Dropout)、深度集成、高效集成近似,以及最后一层或单遍方法。我们还梳理了证据网络、先验网络、共形预测与事后校准等相关工作,以及分布外检测、选择性预测等决策时间任务。对每类方法,我们均考察其理论动机、实现方式、实证性能与局限性。随后,我们综述集成多样性理论与不确定性度量及其分解,对比熵分解与成对散度度量,并整合评估方法,使定性比较基于共同基准。最后,简要探讨大语言模型中的不确定性及开放研究方向,包括分类的认知型高效度量、最后一层多样性、分布偏移下的多样性与校准,以及混合架构。
英文摘要
The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures lack principled uncertainty quantification. This survey provides a structured, critical review of methods for Uncertainty Quantification (UQ) in deep learning, scoped to ensemble-based and approximate Bayesian approaches and the measures used to summarize their outputs. Relative to existing UQ surveys, our contribution is depth on efficient ensemble approximations and single-pass methods, and a unified treatment that separates the method producing a predictive distribution from the measure that summarizes its uncertainty. We organize methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer or single-pass approaches. We situate adjacent work on evidential and prior networks, conformal prediction, and post-hoc calibration, together with the decision-time tasks of out-of-distribution detection and selective prediction. For each, we examine theoretical motivation, implementation, empirical performance, and limitations. We then review ensemble diversity theory and uncertainty measures and their decompositions, contrasting the entropy decomposition with pairwise divergence measures, and consolidate evaluation methodology so that our qualitative comparisons share a common basis. We close with a brief treatment of uncertainty in large language models and open research directions, including efficient epistemic measures for classification, last-layer diversity, diversity and calibration under shift, and hybrid architectures.
发表机构
- Faculty of Computer Science, Dalhousie University(达尔豪斯大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。