发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散模型测试时对齐问题,提出 Best-of-$N$ 引导 (BoNG) 方法,将 BoN 选择融入反向扩散过程,通过粒子间非对称引导提升最终样本及平均质量,在多数比较中表现最优并支持多输出。
AI 中文摘要
扩散模型在生成方面表现出强大的性能,但往往难以使生成的样本与由奖励模型衡量的人类偏好对齐。一种简单而有效的测试时对齐算法是 Best-of-$N$ (BoN) 采样,它从预训练的扩散模型中抽取 $N$ 个独立同分布的样本,并输出奖励最高的单个样本。尽管 BoN 采样在经验上取得了成功,但它对奖励信息的利用有限,因为奖励仅在最终选择阶段被纳入,并未在采样过程中影响反向扩散轨迹。因此,BoN 采样并未提高生成样本的平均对齐程度,并且主要适用于单输出场景。我们提出了 Best-of-$N$ 引导 (BoNG),一种将 BoN 采样原理直接整合到反向扩散过程中的新方法。BoNG 对去噪粒子执行在线 BoN 选择,并调整反向扩散过程,以在生成过程中将粒子群引导至更高奖励区域。具体而言,通过引入去噪粒子之间的非对称引导交互,BoNG 将当前的 BoN 粒子作为引导信号传递给其余粒子群。这种粒子级交互重塑了采样过程,使其朝向更高奖励区域,从而使 BoNG 不仅能在最终最佳样本上超越 Vanilla BoN 采样,还能提高生成样本的平均质量。在 36 项实证比较中,BoNG 在 29 个案例中取得了最佳性能,在 80.56% 的比较中排名第一,优于 SMC 和 Vanilla BoN 采样。BoNG 还支持多输出能力,其 ImageReward 分数达到最新基于样本的引导方法的 1.3 倍,同时实现了 1.6 倍的加速。我们在该 https URL 发布了代码。
英文摘要
Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporated only at the final selection stage without influencing the reverse diffusion trajectory during sampling. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-$N$ Guidance (BoNG), a novel method that integrates the principle of BoN sampling directly into the reverse diffusion process. BoNG performs online BoN selection over denoising particles and adjusts the reverse diffusion process to steer the particle population toward higher-reward regions during generation. Specifically, by introducing an asymmetric guidance interaction among denoising particles, BoNG uses the current BoN particle as a guidance signal to the rest of the particle population. This particle-level interaction reshapes the sampling process toward higher-reward regions, enabling BoNG to improve not only the final best sample beyond Vanilla BoN sampling, but also the average quality of generated samples. Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56% of the comparisons against SMC and Vanilla BoN sampling. BoNG also supports multi-output capability, achieving 1.3$\times$ ImageReward score of the latest sample-based guidance method with a 1.6$\times$ speedup. We release the code at https://github.com/aailab-kaist/BoNG.
CommentsNeurIPS 2026