PosteriorBench:从点估计到后验匹配的生成式逆求解器评估
PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers
另 3 家 · 查看机构详情
- California Institute of Technology(加州理工学院)
- National Taiwan University(台湾大学)
- Lawrence Berkeley National Laboratory(劳伦斯伯克利国家实验室)
- Stanford University(斯坦福大学)
- EarthFlow AI, Inc.(EarthFlow AI公司)
- Imperial College London(伦敦帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对生成式逆求解器仅评估单点重建的不足,提出PosteriorBench基准,通过高保真参考后验与五指标套件评估分布精度,揭示现有求解器的分布匹配差距。
中文摘要 AI 辅助
生成模型越来越多地被用于解决科学逆问题,但现有的评估仍主要关注方法能否产生单一合理的重建结果。这对于不适定问题而言是不够的,因为在不适定问题中,多个解可能与相同的稀疏或噪声观测一致。在这些情况下,一种方法可能实现较强的逐点精度,但仍因模式坍缩、过度自信的不确定性或对不相容解取平均而无法捕获真实后验。我们提出了PosteriorBench,一个用于评估生成式逆求解器分布精度的基准。PosteriorBench评估四个基于物理的逆问题:达西流反演、泊松源恢复、碳捕集与封存以及光传输材料推断。对于每个任务,我们使用计算量大但成熟的方法(如拒绝采样和马尔可夫链蒙特卡洛)构建高保真参考后验,从而能够直接评估求解器是否恢复完整的解集而非单个最佳样本。我们将这些参考与一个包含五个指标的评估套件配对:后验均值误差、后验标准差误差、最大均值差异、切片Wasserstein距离和径向平均功率谱误差。这些指标评估逐点精度、边际不确定性、分布对齐和全局频率保真度。该基准涵盖稀疏感知、低分辨率观测、非线性正演模型、不同噪声水平以及多模态先验,并提供用于分布匹配和不确定性量化的统一流程。我们的实验揭示了当前求解器在分布匹配上的显著差距,同时表明神经算子提高了分辨率鲁棒性,且引导权重和生成噪声是后验方差校准的关键。
英文摘要
Generative models are increasingly used to solve scientific inverse problems, but existing evaluations still focus primarily on whether a method can produce a single plausible reconstruction. This is insufficient for ill-posed problems, where multiple solutions may be consistent with the same sparse or noisy observations. In these settings, a method can achieve strong pointwise accuracy while still failing to capture the true posterior through mode collapse, overconfident uncertainty, or averaging incompatible solutions. We introduce PosteriorBench, a benchmark for evaluating the distributional accuracy of generative inverse solvers. PosteriorBench evaluates four physics-based inverse problems: Darcy flow inversion, Poisson source recovery, carbon capture and storage, and light transport material inference. For each task, we construct high-fidelity reference posteriors using computationally heavy but established procedures such as rejection sampling and Markov chain Monte Carlo, enabling direct assessment of whether solvers recover the full set of solutions rather than the single best sample. We pair these references with a five-metric posterior evaluation suite: posterior-mean error, posterior-standard-deviation error, maximum mean discrepancy, sliced Wasserstein distance, and radially averaged power-spectrum error. These metrics assess pointwise accuracy, marginal uncertainty, distributional alignment, and global frequency fidelity. The benchmark spans sparse sensing, low-resolution observations, nonlinear forward models, varying noise levels, and multimodal priors, with a unified pipeline for distribution matching and uncertainty quantification. Our experiments reveal substantial distribution-matching gaps across current solvers, while showing that neural operators improve resolution robustness, and guidance weights and generation noise are key to posterior-variance calibration.