发表机构
CWI; University of Twente; University College London; Advanced Research Computing Centre, University College London(荷兰国家数学和计算机科学研究院; 特文特大学; 伦敦大学学院; 伦敦大学学院高级研究计算中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对神经网络全贝叶斯推断UQ计算不可行的问题,提出广义拉普拉斯主动子空间方法,通过均值损失缩放得到稳定低维子空间,实现可扩展的校准UQ。
AI 中文摘要
可靠的不确定性量化(UQ)是将神经网络部署到科学领域及高风险应用场景的关键,但对网络参数进行全贝叶斯推断在计算上不可行。我们提出一种基于少量数据驱动曲率方向的低秩广义拉普拉斯近似,用于神经网络UQ。从通过经验损失定义的广义贝叶斯后验出发,我们在预训练权重集合的该主动曲率子空间内构造局部高斯近似。保留子空间内的后验方差可通过闭式形式得到,且先验方差通过经验贝叶斯过程校准。该广义贝叶斯公式允许我们比较两种后验缩放方式:与求和负对数似然相关的标准贝叶斯缩放,以及将经验损失按数据数量归一化的均值损失缩放。核心发现是,标准缩放会导致主导主动方向上的后验方差出现依赖数据量的收缩。在回归问题中,这可能迫使低秩框架保留额外的弱曲率方向,以实现校准数据的名义覆盖率。当后验样本通过非线性网络传播时,这些额外方向会降低预测区间的一致性,并使后验预测均值偏离预训练模型。相比之下,广义均值损失缩放会产生更稳定、更低维的主动子空间,并生成校准后的、一致的预测置信区间。这些结果表明,广义拉普拉斯主动子空间为神经网络中校准不确定性量化提供了一种实用且可扩展的途径。
英文摘要
Reliable uncertainty quantification (UQ) is essential for deploying neural networks in scientific and high-stakes applications, but full Bayesian inference over the network parameters is computationally infeasible. We propose a low-rank generalized Laplace approximation for neural-network UQ based on a small number of data-informed curvature directions. Starting from a generalized Bayesian posterior defined through an empirical loss, we construct a local Gaussian approximation around a pretrained set of weights in this active curvature subspace. The posterior variances in the retained subspace are available in closed form, and the prior variance is calibrated by an empirical Bayes procedure. The generalized Bayesian formulation allows us to compare two posterior scalings: the standard Bayesian scaling associated with the summed negative log likelihood, and a mean-loss scaling in which the empirical loss is normalized by the number of data. A central finding is that the standard scaling induces a data-size dependent contraction of the posterior variance in the leading active directions. In regression problems, this can force the low-rank framework to retain additional weak-curvature directions in order to achieve nominal coverage of calibration data. When posterior samples are propagated through the non-linear network, these additional directions can degrade the coherence of the predictive intervals and shift the posterior predictive mean away from the pretrained model. In contrast, the generalized mean-loss scaling yields a more stable, lower dimensional active subspace and produces calibrated, coherent predictive confidence intervals. These results indicate that generalized Laplace active subspaces provide a practical and scalable route to calibrated uncertainty quantification in neural networks.