arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于自然语言条件生成先验用于偏微分方程中的贝叶斯反演

Generative Priors Conditioned on Natural Language for Bayesian Inversion in PDEs

Pengyu Zhang, Mark Girolami, Arnaud Vadeboncoeur

arXiv 2609.32941首次发表:更新:

AI 中文总结

本文提出利用自然语言条件生成先验,结合定量与定性数据,对偏微分方程中的感兴趣量进行贝叶斯反演与不确定性量化,通过条件扩散和条件自编码器两种方法实现,并在多个物理系统上验证。

AI 中文摘要

从数据中推断感兴趣量(QoI)是科学与工程中的核心任务。在此类情境中,我们通常同时获得定量数据和定性数据。定量数据可表示为含噪声的传感器测量、模拟数据、再分析数据;定性数据则可能以实验设置的文字描述、预期实验结果以及人类感知的系统行为等形式存在。本文所解决的问题如下:给定一组配对的定性文本和定量QoI数据训练集,我们学习利用两种模态之间的内在相关性,以学习一个信息丰富的数据驱动的自然语言条件贝叶斯先验,使得当面对一个新的物理系统时,我们能够连贯地结合(i)训练数据集,(ii)描述新系统的定性文本,以及(iii)来自该新系统的少量含噪声传感器读数,对QoI进行推断和不确定性量化(UQ)。为实现此任务,我们开发了两种并行方法,一种使用条件扩散,另一种使用条件自编码器,并将两者与经典贝叶斯方法、无条件生成模型和确定性监督方法进行比较。每种方法都有特定的优势和权衡;条件自编码器提供理论可处理性,允许快速后验采样,并提供更校准的UQ,而条件扩散则被探索用于更强的表达能力和捕获具有不规则QoI场的复杂后验。该方法在稳态热方程、阻尼亥姆霍兹方程和英国天气再分析数据上进行了测试。

英文摘要

Inferring quantities of interest (QoI) from data is a central task in Science and Engineering. In such contexts, we often have access to both quantitative data and qualitative data. Quantitative data may be represented by noisy sensor measurements, simulation data, re-analysis data; qualitative data may be in the form of text descriptions of experimental setups, expected experiment outcomes, and human-perceived system behaviours. The task we address in this paper is the following. Given a training set of paired qualitative text and quantitative QoI data, we learn to exploit the inherent correlation between the two modalities to learn a highly informative data-driven natural-language-conditional Bayesian prior, such that when presented with a new physical system, we can coherently combine (i) the training dataset, (ii) qualitative text describing the new system, and (iii) a small number of noisy sensor readings from that new system, to perform inference and uncertainty quantification (UQ) over the QoI. To achieve this task, we develop two parallel approaches, one uses conditional diffusion and the other conditional autoencoders, and compare both against classical Bayesian methodology, unconditional generative models and deterministic supervised methods. Each approach has specific strengths and tradeoffs; conditional autoencoder offers theoretical tractability, allows for fast posterior sampling, and provides better-calibrated UQ, whereas conditional diffusion is explored for greater expressiveness and capturing complex posteriors with irregular QoI fields. The approach is tested on the steady-state heat equation, damped Helmholtz equation, and UK weather reanalysis data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑