发表机构
University of Cambridge; Spotify(剑桥大学; Spotify)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对拉普拉斯近似中先验精度选择对大神经网络代价过高的问题,基于最大更新参数化推导先验协方差重标度,实现从小型模型到大型模型的零样本不确定性迁移,大幅提升精度搜索速度且预测性能损失极小。
AI 中文摘要
拉普拉斯近似中可靠的预测不确定性高度依赖先验精度,但选择该精度需进行后验搜索,这对拥有数十亿参数的神经网络而言代价过高。在最大更新参数化($\mu\mathrm{P}$)下,我们推导了先验协方差的重标度方法,使所选精度随模型宽度增长保持稳定。这引出了$\sigma\mathrm{Transfer}$:我们在较小模型上选择精度,再将其零样本迁移至大得多的模型,即完全无需在大模型上搜索精度。我们证明了先验核、后验协方差、所选精度及后验导出决策在明确条件下的收敛性,并在回归、图像分类和Transformer读出任务中验证了$\sigma\mathrm{Transfer}$。例如,在MNIST上从宽度128迁移至4096时,测得的精度搜索加速比达$\sim 5000\times$,目标负对数似然(NLL)下降为$0.002$;从公开的1B模型迁移至7B模型时,10项任务的中位数搜索加速比为$\sim 2.3\times$(最高达$\sim 330\times$),平均测得的目标NLL增幅低于$10^{-4}$。相同的后验稳定性还支持迁移获取函数、分布外检测及弃权决策,无需构建目标后验。
英文摘要
Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters. Under the Maximal Update Parametrization ($μ\mathrm{P}$), we derive a rescaling of the prior covariance that makes the selected precision stable as model width grows. This leads to $σ\mathrm{Transfer}$: we select the precision on a smaller model and zero-shot transfer it to the much larger model, i.e., without searching for the precision on the larger model at all. We show convergence of the prior kernel, posterior covariance, selected precision, and posterior-derived decisions under explicit conditions, and verify $σ\mathrm{Transfer}$ across regression, image classification, and Transformer readouts. For example, measured precision-sweep speedups reach $\sim 5000\times$ when transferring from width 128 to 4096 on MNIST, at a target-NLL degradation of $0.002$; transferring from a public 1B to 7B model gives a median search speedup of $\sim 2.3\times$ (up to $\sim 330\times$), with a mean measured target-NLL increase below $10^{-4}$ across ten tasks. The same posterior stability also enables transfer of acquisition, OOD-detection, and abstention decisions without constructing a target posterior.