发表机构
University of Tübingen; Hertie Institute for AI in Brain Health; Tübingen AI Center; Charité–Universitätsmedizin Berlin; Bernstein Center for Computational Neuroscience Berlin(蒂宾根大学; 赫蒂脑健康人工智能研究所; 蒂宾根人工智能中心; 柏林夏里特医学院; 柏林伯恩斯坦计算神经科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出将不确定性统一定义为逐点后验风险,构建了可直接计算神谕不确定性的理论基准,发现准确预测不保证可靠的不确定性解耦,该基准可分析不同方法的不确定性估计性能。
AI 中文摘要
可靠的不确定性估计对安全敏感应用至关重要,理解预测不确定性的来源是关键,这通常需要将认知不确定性(epistemic uncertainty)与偶然不确定性(aleatoric uncertainty)解耦,但文献中对这些不确定性类型的定义并不一致,难以评估方法是否能产生准确的不确定性估计。由于真实的认知不确定性通常不可用,评估进一步复杂化,现有基准大多依赖分布外检测等代理任务,这些任务无法提供完整的真实不确定性目标,对不确定性估计的结构和质量的洞察有限。我们提出将不确定性统一定义为逐点后验风险,即给定数据时,预测器在合理真实函数分布下的期望损失。该视角结合了函数上的贝叶斯不确定性与估计量相对于后验均值的偏差,捕捉模型误设、优化误差等效应。该公式构成了基于理论的基准的基础,该基准可利用具有真实协变量和已知生成过程的半合成数据集直接计算神谕(oracle)认知和偶然不确定性,通过避免代理评估,该基准实现对不确定性估计的细粒度分析。我们的实证研究发现,准确的预测并不能保证可靠的不确定性解耦;该基准揭示了不同方法间具有实际意义的差异,识别出与神谕不确定性目标有合理一致性的方法,同时暴露了这些方法对数据集和建模选择的敏感性。
英文摘要
Reliable uncertainty estimates are critical in safety-sensitive applications. For such estimates to be useful in practice, it is crucial to understand the sources underlying a model's uncertainty, motivating the disentanglement of total uncertainty into epistemic and aleatoric uncertainty. Existing notions of uncertainty differ in the sources they capture and, consequently, in their definitions of aleatoric and epistemic uncertainty, with no universally accepted definition. We define uncertainty through sample-conditional pointwise posterior risk, which is the expected loss of a predictor under the distribution of plausible ground-truth functions given the observed sample. This definition unifies probabilistic and risk-based concepts of uncertainty. To assess state-of-the-art uncertainty disentanglement methods, we develop a framework that directly compares their estimates against ground-truth uncertainty defined primarily by posterior risk, alongside commonly used alternative uncertainty definitions. We find that Spectral-normalized Neural Gaussian Processes and Variational Latent Gaussian Processes most closely recover the ground-truth uncertainty, while most methods track posterior variance more closely than posterior risk, missing the predictors' bias. Beyond method rankings, we investigate how strongly estimated aleatoric and epistemic uncertainty are entangled and how sensitive uncertainty quality is to modeling choices, yielding practical guidance for uncertainty disentanglement. To support further method development and validation, we release 13 semi-synthetic UCI/OpenML datasets with known posteriors, enabling the computation of ground-truth uncertainty. Code and data will be made publicly available upon acceptance.