发表机构
Uppsala University(乌普萨拉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出序列控制交互式多粒子流图(IMPFM),通过多粒子交互和流图驱动的后验样本共享机制,实现在线反馈驱动搜索中的全局探索与偏好对齐,避免模式崩溃和权重退化。
AI 中文摘要
虽然生成模型已经实现了无需训练的奖励对齐,但当前方法通常擅长在底层分布的狭窄区域内进行局部探索。当偏好先验未知且仅通过序列反馈揭示时,这些方法难以应对——这种情况需要广泛探索以发现高效用区域。为了解决这个问题,我们提出了序列控制交互式多粒子流图(IMPFM),一个样本高效的在线反馈驱动搜索框架。IMPFM逐步将一组交互粒子向目标分布传输,保持异质偏好对齐所需的广泛覆盖。IMPFM引入了一种基于流图的原则性且高效的后验样本共享机制。通过在每次重采样步骤中用整个集成体的集体后验样本纠正单个粒子漂移,该框架最大化样本效用以实现全局探索,同时主动缓解标准控制框架中常见的奖励过度优化问题。结合涉及多粒子交互的原则性探索-利用重新加权机制,这种序列校正的多粒子动力学明确保留了结构多样性,并克服了标准SMC采样器固有的权重退化问题。关键的是,我们证明了所得到的采样框架产生了一个多粒子交互感知的Feynman-Kac校正器,逐步将多粒子系统引导向KL倾斜的目标分布,促进全局探索并防止模式崩溃。在多种搜索和对齐任务上的广泛经验评估和严格消融实验证实了IMPFM相对于现有基线的有效性。
英文摘要
While generative models have enabled training-free reward alignment, existing particle-based methods are fundamentally constrained by their propensity to local exploration within narrow regions of the underlying distribution, severely restricting sample diversity. This limitation becomes especially acute under tight reward-feedback budgets, where effective search demands broad, strategic exploration to uncover high-utility regions. To address this, we propose Sequentially-Controlled Interactive Multi-Particle Flow-Maps (IMPFM), a framework for feedback-efficient search. IMPFM progressively transports a group of interactive particles toward the target distribution, maintaining the broad coverage essential for heterogeneous preference alignment. IMPFM leverages a principled and efficient posterior sample-sharing mechanism across particles powered by flow maps. By correcting individual particle drift with the collective value gradient from the entire ensemble's posterior samples at each correction step, the framework maximizes sample utility to enable global exploration while actively mitigating reward over-optimization, typical of standard control frameworks. Paired with a principled exploration-exploitation reweighting mechanism involving multi-particle interaction, this sequentially corrected multi-particle dynamics explicitly preserves structural diversity and overcomes the weight degeneracy inherent to standard Sequential Monte Carlo (SMC) samplers. Crucially, we prove that the resulting sampling framework yields a multi-particle interaction-aware Feynman-Kac corrector that progressively steers the multi-particle system toward a KL-tilted target distribution, facilitating global exploration and preventing mode collapse. Extensive empirical evaluations and ablations across diverse search and alignment tasks confirm the efficacy of IMPFM over existing baselines.
Comments27 pages, 20 figures; Preprint