发表机构
Université Grenoble Alpes; KAUST(格勒诺布尔阿尔卑斯大学; 阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出通过深度权重分解先验尺度来学习贝叶斯神经网络的随机参数划分,以稀疏化随机性而非容量,在保持性能的同时减少随机参数数量。
AI 中文摘要
贝叶斯神经网络无需完全随机即可成为通用条件密度逼近器,但哪些参数应该是随机的仍然是一个开放问题。我们通过将深度权重分解应用于先验尺度(即参数先验的标准差)来学习这种划分,同时使用最大均值差异目标将函数先验拟合到高斯过程。先验尺度低于阈值的参数变为确定性参数,并在推理过程中被优化,因此正则化器稀疏化的是随机性而非容量。我们给出了一个可在线性时间内检查的通用条件密度逼近的证书,并在证书失败时提供最小修复。我们进一步证明,常见的混合方案(采样部分参数并优化其余参数)是类型II最大后验目标的随机近似,并且耦合步长可能留下一个不随步长缩小而消失的跟踪误差。在一个双峰目标上,学习到的划分在所有预算下都接近无约束参考,并且对阈值不敏感,而将相同先验尺度随机分配到各层的随机掩码则最多差两个数量级。在UCI基准上,我们的方法与完全随机网络性能相当,同时保持约一半参数为确定性参数。
英文摘要
Bayesian neural networks need not be fully stochastic to be universal conditional density approximators, but it remains open which parameters should be stochastic. We learn this split by applying deep weight factorization to the prior scales, which are the standard deviations of the parameter priors, while fitting the functional prior to a Gaussian process with a maximum mean discrepancy objective. A parameter whose prior scale falls below a cutoff becomes deterministic and is optimized during inference, so the regularizer sparsifies stochasticity rather than capacity. We give a certificate for universal conditional density approximation that is checkable in linear time, together with a minimal repair when it fails. We further show that the common hybrid scheme of sampling some parameters and optimizing the others is stochastic approximation for a type-II maximum a posteriori objective, and that coupled step sizes can leave a tracking error that does not vanish as the step size shrinks. On a bimodal target, the learned split stays close to an unconstrained reference across all budgets and is insensitive to the cutoff, while random masks that distribute the same prior scales across layers are worse by up to two orders of magnitude. On UCI benchmarks, our method performs on par with a fully stochastic network while keeping about half of its parameters deterministic.