利用昂贵模拟器贝叶斯推理中的梯度信息
Exploiting Gradients in Bayesian Inference of Expensive Simulators
浏览论文内容
中文总结 AI 辅助
本文将梯度信息融入高斯过程代理,在有限模拟预算下加速昂贵模拟器的贝叶斯推理,反向模式微分下效率提升显著,适用于参数数超输出维度的场景。
中文摘要 AI 辅助
基于微分方程的模拟器在科学与工程领域中无处不在,它们常被用于基于模拟的推理,以根据模拟器输出的真实观测值评估输入参数的后验分布。然而,当单个模拟器评估的计算成本很高时,推理会变得极具挑战性。在这种情况下,人们采用基于贝叶斯优化的主动学习方法,结合高斯过程代理模型,以在有限的模拟预算下最大化获取的信息。近年来,模拟器输出相对于输入参数的梯度信息越来越容易获取,但却很少被用于推理。尽管我们仅需学习模拟器的输入-输出关系,梯度信息仍能提供额外的有价值信号来指导主动学习过程,这在样本效率至关重要的昂贵模拟器场景中尤为重要。本文展示了如何将梯度信息融入高斯过程代理,以在有限的模拟预算下加速基于贝叶斯优化的推理。结果表明,使用梯度信息可显著提升收敛速度;对于反向模式微分,在考虑额外计算成本的情况下,推理效率的提升仍能保持;而对于前向模式微分,推理速度的提升无法抵消计算成本。这些结果表明,梯度增强代理主要在参数数量超过输出维度且反向模式微分高效的问题中具有优势。
英文摘要
Simulators based on differential equations are ubiquitous in science and engineering. They are often used in simulation-based inference to evaluate the posterior distribution of the input parameters based on real-world observations of the simulator outputs. However, inference becomes challenging when individual simulator evaluations are computationally expensive. In such cases, a Bayesian optimization-based active learning approach with Gaussian process surrogate models has been used to maximize the information obtained from a limited simulation budget. Recently, gradients of simulator outputs with respect to input parameters have become increasingly available, yet they are rarely exploited for inference. Even though we only need to learn the simulator input-output relationship, gradient information can provide an additional valuable signal to guide the active learning procedure. This is of particular interest in the case of expensive simulators, when sample efficiency is crucial. In this paper, we demonstrate how incorporating gradient information into the Gaussian process surrogate accelerates Bayesian optimization-based inference under a limited simulation budget. Our results show significant improvement in convergence speed from using gradient information. For reverse-mode differentiation, the inference efficiency gains are maintained when accounting for the additional computational cost. In contrast, for forward-mode differentiation, the inference speed-up does not outweigh the computational costs. These results indicate that gradient-enhanced surrogates are beneficial primarily in problems where the number of parameters exceeds the output dimensionality, where reverse-mode differentiation is efficient.
发表机构
- Czech Technical University in Prague(布拉格捷克技术大学)
- Institute of Plasma Physics of the Czech Academy of Sciences(捷克科学院等离子体物理研究所)
机构由 AI 辅助整理,请以论文原文为准。