AI 中文总结
本文提出一种主动学习框架,平衡探索与复制以实现数据高效的随机仿真模型校准,经实验验证可减少仿真数量并提升后验分布学习效果。
AI 中文摘要
基于仿真的校准旨在通过使模型输出与现实世界观测结果对齐,来推断复杂仿真模型的未知参数。当仿真运行计算成本高昂时,可使用基于仿真数据训练的统计模拟器来高效近似该模型。为构建模拟器而进行的智能自适应仿真输入选择,可大幅提升校准过程的效率。对于带有噪声输出的随机仿真,该任务极具挑战性,因为选择新输入位置(探索)和在现有输入处分配重复运行(复制),对于高效学习输入-输出关系都至关重要。本文提出一种主动学习框架,该框架可自适应平衡探索与复制,以实现数据高效校准。我们的不确定性感知采集准则旨在学习未知仿真参数的后验密度,并推导了两种对应的采集函数形式,分别用于探索和复制。基于此,我们提出一种策略,在序列设计的每个阶段,在探索和复制之间进行选择,以最有效地降低仿真参数后验密度估计的不确定性。在合成基准测试和真实流行病学模型上开展的实验表明,我们的方法可显著提升仿真参数后验分布的学习效果,同时减少所需仿真的数量,因此非常适用于昂贵的随机仿真场景。
英文摘要
Simulation-based calibration aims to infer unknown parameters of complex simulation models by aligning model outputs with real-world observations. When simulation runs are computationally expensive, statistical emulators trained on simulation data are used to efficiently approximate the model. An intelligent, adaptive selection of simulation inputs for building the emulator can substantially improve the efficiency of the calibration process. This task is particularly challenging for stochastic simulations with noisy outputs, since both selecting new input locations (exploration) and allocating repeated runs at existing inputs (replication) are essential for efficiently learning the input-output relationship. In this paper, we introduce an active learning framework that adaptively balances exploration and replication for data-efficient calibration. Our uncertainty-aware acquisition criterion targets learning the posterior density of the unknown simulation parameters, and we derive two corresponding forms of the acquisition function for exploration and replication. Building on these, we propose a strategy that, at each stage of the sequential design, chooses between exploration and replication to most effectively reduce the uncertainty in the estimate of the posterior density of the simulation parameters. Experiments on synthetic benchmarks and a real epidemiological model demonstrate that our approach significantly improves learning of the posterior distribution of the simulation parameters while reducing the number of required simulations, making it well-suited for expensive stochastic simulation settings.
DOI:10.1080/00224065.2026.2710639