arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不完美扩散模型的纠错推理时扩展

Error-Corrected Inference-Time Scaling for Imperfect Diffusion Models

Zuokai Wen, Louis Grenioux, Weinan E, Jiequn Han

arXiv 2610.01933首次发表:更新:

发表机构

Shanghai Jiao Tong University; Flatiron Institute; Peking University(上海交通大学; 熨斗研究院; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对不完美扩散模型,提出基于能量的费曼-卡克校正器(EBFKC),通过顺序蒙特卡洛和能量代理校正推理时采样误差,实验验证其能紧密匹配目标分布。

AI 中文摘要

推理时扩展使预训练扩散模型无需额外训练即可适应新的采样任务。现有方法主要依赖使用更多粒子的蒙特卡洛采样,但前提是预训练模型是精确的。实际上,数据和训练限制使模型不完美,这些方法继承了其误差。更多粒子减少了蒙特卡洛误差,但无法消除端点与目标之间的失配或跟踪规定概率路径时的误差。我们引入了基于能量的费曼-卡克校正器(EBFKC),这是一个针对基于能量的扩散模型的框架,在给定参考能量的情况下,能够即时纠正这些误差。我们首先推导了费曼-卡克动力学,即使在模型不完美的情况下,也能在连续时间总体极限中精确跟踪规定路径,并使用带有方差控制引导的顺序蒙特卡洛方法近似这些动力学。为了消除端点失配,我们沿扩散路径使用预训练能量作为代理,并逐步纳入学习到的终端能量与目标终端能量之间的差异。在高斯混合模型、粒子系统、丙氨酸二肽和丙氨酸四肽上的实验表明,在退火和奖励倾斜下,我们的方法能紧密匹配目标分布和分子自由能分布,而标准的推理时扩展基线则保留了显著的采样误差。

英文摘要

Inference-time scaling adapts pretrained diffusion models to new sampling tasks without additional training. Existing methods rely primarily on Monte Carlo sampling with more particles, yet are premised on the pretrained model being exact. In practice, data and training limitations make the model imperfect, and these methods inherit its error. More particles reduce Monte Carlo error but cannot remove the mismatch between the endpoint and the desired target or the error in tracking the prescribed probability path. We introduce the Energy-based Feynman-Kac Corrector (EBFKC), a framework for energy-based diffusion models that corrects these errors on the fly given a reference energy. We first derive Feynman-Kac dynamics that track a prescribed path exactly in the continuous-time population limit even when the model is imperfect, and approximate these dynamics using sequential Monte Carlo with variance-controlling guidance. To remove the endpoint mismatch, we use the pretrained energy as a surrogate along the diffusion path and progressively incorporate the discrepancy between the learned and target terminal energies. Experiments on Gaussian mixture models, particle systems, alanine dipeptide, and alanine tetrapeptide show that our method closely matches target distributions and molecular free-energy profiles under annealing and reward tilting, whereas standard inference-time scaling baselines retain substantial sampling errors.

CommentsUnder review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑