arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

似然温度调节对变分贝叶斯线性神经网络极限预测矩的影响

The Impact of Likelihood Tempering on the Limiting Predictive Moments of Variational Bayesian Linear Neural Networks

Ian Zhang, Thibault Randrianarisoa

arXiv 2610.09132首次发表:更新:

发表机构

University of Toronto; Dunlap Institute for Astronomy and Astrophysics(多伦多大学; 邓拉普天文学与天体物理研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究变分贝叶斯线性神经网络中似然温度调节对极限预测矩的影响,发现预测期望和方差在不同温度衰减尺度下发生相变,并可通过调节温度参数恢复NNGP后验的期望或方差。

AI 中文摘要

在宽贝叶斯神经网络中,高斯均值场变分推断容易出现“先验主导”问题:ELBO中的Kullback-Leibler (KL)正则项超过期望对数似然,随着宽度$M$增大,变分预测分布坍缩为先验预测。通过将似然提高到$1/T$次幂(温度$T<1$)进行温度调节,等价于将KL项缩放$T$倍。本文研究$T$必须随$M$以多快速度减小以抵消这种退化,并在两项之间取得良好平衡。对于具有各向同性高斯先验的单隐层线性网络,我们推导了在$T = \tau/M^{c}$(常数$\tau, c > 0$)形式调度下,当$M \to \infty$时的极限预测分布,并将其与未调节温度的神经网络高斯过程(NNGP)后验(精确后验的无限宽度极限)进行比较。我们的主要结果是预测期望和方差在不同尺度上发生相变:极限期望在$c = 1/2$时离开其先验值,一旦$\tau$低于显式阈值,且对于$c > 1/2$等于最小二乘预测,而极限方差在$c < 1$时保持其先验值,在$c=1$时匹配NNGP的方差,在$c > 1$时消失。通过适当选择$\tau,c$,可以恢复NNGP后验期望或其方差。

英文摘要

In wide Bayesian neural networks, Gaussian mean-field variational inference is prone to "prior dominance": the Kullback-Leibler (KL) regularization term of the ELBO outweighs the expected log-likelihood, and the variational predictive distribution collapses to the prior predictive as the width $M$ grows. Tempering the likelihood, by raising it to the power $1/T$ for a temperature $T < 1$, is equivalent to scaling the KL term by $T$. We ask in this paper how fast $T$ must decrease with $M$ to counteract this degeneracy and strike a good balance between the two terms. For single-hidden-layer linear networks with isotropic Gaussian priors, we derive the limiting predictive distribution under schedules of the form $T = τ/M^{c}$, with constants $τ, c > 0$, as $M \to \infty$ and compare it with the untempered neural network Gaussian process (NNGP) posterior, the infinite-width limit of the exact posterior. Our main result is that the predictive expectation and variance undergo phase transitions at different scales: the limiting expectation leaves its prior value at $c = 1/2$, once $τ$ falls below an explicit threshold, and equals the least-squares prediction for $c > 1/2$, whereas the limiting variance keeps its prior value for $c < 1$, matches the NNGP's for $c=1$, and vanishes for $c > 1$. With suitable choices of $τ,c$, one can recover either the NNGP posterior expectation or its variance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑