arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用优化器派生几何引导后验探索

Guiding Posterior Exploration with Optimizer-Derived Geometry

Moritz Schlager, Emanuel Sommer, Thomas Möllenhoff, David Rügamer

arXiv 2607.25312首次发表:更新:

发表机构

TUM; LMU Munich; RIKEN AIP(慕尼黑工业大学; 慕尼黑大学; 理化学研究所先进智能项目中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究贝叶斯神经网络不确定性量化中高维多峰后验分布探索的计算成本问题,提出基于优化器派生几何的预处理采样策略,可减少采样预热阶段,提升性能与稳定性,且无额外计算成本,经多数据集和架构验证。

AI 中文摘要

基于采样的方法为贝叶斯神经网络中的不确定性量化提供了原则性方法。然而,探索高维多峰后验分布的计算成本常常给其实际应用带来挑战。为克服这些困难,贝叶斯深度集成(即从多个优化解热启动采样)已被证明是一种有效策略。本文表明,在自适应优化器(如AdamW)热启动期间作为副产品计算出的曲率估计,能以可忽略的额外成本为采样阶段提供信息。具体而言,我们基于优化器派生几何提出的预处理采样策略,可大幅减少甚至消除冗长的采样预热阶段需求,并带来更高的数值稳定性。该方法在无任何额外计算成本的情况下,持续保持或提升预测性能和不确定性量化。我们在各种数据集和网络架构上证实了研究结果的一致性。

英文摘要

Sampling-based methods offer a principled approach to uncertainty quantification in Bayesian neural networks. Their practical use, however, is often challenged by the computational cost of exploring high-dimensional and multimodal posterior distributions. To overcome these difficulties, Bayesian Deep Ensembles, i.e., warmstarting the sampling from several optimized solutions, have proven to be an effective strategy. In this paper, we demonstrate that curvature estimates computed during the warmstart as a byproduct in adaptive optimizers such as AdamW can inform the sampling phase at negligible additional cost. Specifically, our proposed preconditioned sampling strategy based on optimizer-derived geometries can substantially reduce or even eliminate the need for a lengthy sampling burn-in phase and leads to greater numerical stability. This approach consistently maintains or improves predictive performance and uncertainty quantification without any additional computational costs. We confirm the consistency of our findings across various datasets and network architectures.

CommentsAccepted for presentation at the OPTIMAL Workshop at AISTATS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑