arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可扩展扩散SBI用于模拟器误设下的组合推断

Scalable Diffusion SBI for Compositional Inference under Simulator Misspecification

Vincent D. Zaballa, Elliot E. Hui

arXiv 2609.36950首次发表:更新:

发表机构

University of California, Irvine(加利福尼亚大学尔湾分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对模拟器误设下的组合推断,提出可扩展扩散SBI方法,包括组合采样、分层块状采样和路径正则化微调,并在基准和生物模型上验证了有效性。

AI 中文摘要

当必须组合许多异质观测、必须保留分层潜在结构,且模拟器相对于观测数据被误设时,基于模拟的推断具有挑战性。我们开发了在设计条件设置中用于基于扩散的推断的采样和微调方法,其中同一模拟器在不同实验条件$\xi$下被查询。我们通过一个考虑观测数量的连续时间扩散系数扩展了基于组合分数的推断,避免了雅可比和辅助协方差修正。我们引入了分层块状扩散采样(HBDS),它使用单个预训练模型推断共享参数和组特定潜在状态,层次结构仅在采样时指定。这些方法共同支持可变的观测集和分组,而无需重新训练。为了解决误设问题,我们引入了路径正则化微调,将学习到的似然适应观测,并将修正转移到后验推断。利用Girsanov定理,我们量化了预训练和微调模型在不同实验设计下的路径发散,并将其与预测误差一起解释,以区分候选误设修正与不必要的适应。我们在精确分数高斯和简单似然、复杂后验基准上评估组合采样,在受控分层模型上使用解析和学习的分数评估HBDS,并在一个具有已知设计相关差异的单独解析模型上评估微调和定位。最后,我们将该框架应用于机制性骨形态发生蛋白信号模型中四个细胞系的940个测量值,其中微调相对于预训练模型提高了后验预测准确性,并将后验边缘向最小二乘参考移动,同时保留扩散。

英文摘要

Simulation-based inference is challenging when many heterogeneous observations must be composed, hierarchical latent structure must be preserved, and the simulator is misspecified relative to observed data. We develop sampling and fine-tuning methods for diffusion-based inference in design-conditional settings, where the same simulator is queried across different experimental conditions $ξ$. We extend compositional score-based inference with a continuous-time diffusion coefficient that accounts for the number of observations, avoiding Jacobian and auxiliary-covariance corrections. We introduce Hierarchical Blockwise Diffusion Sampling (HBDS), which infers shared parameters and group-specific latent states using a single pretrained model, with the hierarchy specified only at sampling time. Together, these methods support variable observation sets and groupings without retraining. To address misspecification, we introduce path-regularized fine-tuning that adapts the learned likelihood to observations and transfers corrections to posterior inference. Using Girsanov's theorem, we quantify path divergence between pretrained and fine-tuned models across experimental designs and interpret it alongside predictive errors to distinguish candidate misspecification correction from unnecessary adaptation. We evaluate compositional sampling on exact-score Gaussian and Simple Likelihood, Complex Posterior benchmarks, HBDS with analytic and learned scores on a controlled hierarchical model, and fine-tuning and localization on a separate analytic model with known design-dependent discrepancy. Finally, we apply the framework to 940 measurements across four cell lines in a mechanistic Bone Morphogenetic Protein signaling model, where fine-tuning improves posterior-predictive accuracy relative to the pretrained model and shifts posterior marginals toward the least-squares reference while retaining spread.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑