发表机构
University of Oxford; King's College London(牛津大学; 伦敦国王学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种平均场框架,将扩散模型推理时的分布控制转化为瞄准倾斜测度,推导加权交互粒子方案,为现有批量引导方法提供理论基础,并在低维及蛋白质构象任务中验证了其有效性。
AI 中文摘要
扩散模型正越来越多地被用作可控采样器,其生成结果可在推理阶段根据选定的奖励函数进行引导。这类奖励通常定义在单个样本上,但在许多应用中,更希望根据分布层面的奖励进行引导,例如利用总体层面的信息进行校准或鼓励多样性。在这两种情况下,直接将奖励梯度纳入动力学过程虽常有效,但对采样分布几乎没有理论保证。对于逐点奖励,近期研究已尝试开发基于粒子重加权的原则性框架,以瞄准规定的倾斜分布。然而,目前缺乏针对分布奖励的类似理论基础方法。本研究将推理时的分布控制表述为在平均场框架下瞄准倾斜测度,并推导了一种加权交互粒子方案以原则性地实现该目标。该框架将逐点奖励引导作为特例,同时为现有的批量引导方法提供理论基础。实证方面,我们验证了该过程在易处理的低维设置中能正确瞄准规定分布,并研究了其在高维蛋白质构象任务中的表现。
英文摘要
Diffusion models are increasingly used as controllable samplers, whose generations can be steered at inference time according to a chosen reward function. While such rewards are typically defined on individual samples, for many applications it is desirable to steer according to distribution-level rewards, for example to calibrate with population-level information or to encourage diversity. In both cases, simply incorporating the reward gradient into the dynamics, while often effective, comes with few theoretical guarantees on the sampled distribution. For pointwise rewards, recent work has therefore sought to develop a principled framework for targeting a prescribed tilted distribution using particle reweighting. However, an analogous theoretically-grounded approach for distributional rewards is currently lacking. In this work, we formulate inference-time distributional control as targeting a tilted measure under a mean-field framework, and derive a weighted interacting particle scheme to target it in a principled manner. Our framework recovers pointwise-reward steering as a special case, while providing a theoretical foundation for existing batch-level steering methods. Empirically, we verify that the procedure correctly targets the prescribed distribution in tractable low-dimensional settings, and investigate its behaviour in higher-dimensional protein conformation tasks.
CommentsPreprint. Presented at SPIGM workshop, ICML 2026