发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ZeNOVA,一种无梯度的初始噪声优化方法,通过退火软值引导、流形约束超球朗之万动力学和Metropolis-Hastings跳跃,在图像和视频生成模型中实现稳定高效的奖励对齐。
AI 中文摘要
近期在蒸馏和流图模型方面的进展使得高质量数据的确定性一步或几步生成成为可能,从而促进了奖励对齐方法的一个新分支,该分支直接优化来自高斯分布的初始噪声。然而,大多数现有的初始噪声优化方法依赖于一阶梯度信息,这在黑盒奖励场景中要么不适用,要么存在不稳定和低效的问题。在此,我们介绍了ZeNOVA,一种以无梯度方式进行的稳定且高效的初始噪声对齐方法。具体来说,我们通过退火软值引导、流形约束的超球朗之万动力学以及Metropolis-Hastings跳跃来解决现有算法在黑盒场景中的主要挑战。在图像和视频生成模型上的大量实验表明,ZeNOVA在优化初始噪声以获得更高奖励方面显著更稳定,同时利用高斯先验的几何结构,超越了所有评估的零阶基线方法,展示了其对各种黑盒奖励对齐的实际适用性。
英文摘要
Recent advances in distillation and flow-map models have enabled deterministic one- or few-step generation for high-quality data, facilitating a new branch of reward alignment approaches that directly optimize the initial noise from a Gaussian distribution. However, most existing initial-noise optimization methods rely on first-order gradient information, which is either inapplicable or suffers from instability and inefficiency in black-box reward scenarios. Here, we introduce ZeNOVA, a stable and efficient initial noise alignment method in a gradient-free manner. Specifically, we address existing algorithms' major challenge in black-box scenarios through annealed soft-value guidance, manifold-constrained hyperspherical Langevin dynamics, and Metropolis-Hastings jumping. Extensive experiments on image and video generative models show that ZeNOVA outperforms all evaluated zeroth-order baselines by optimizing the initial noise toward higher rewards substantially more stably while exploiting the geometry of the Gaussian prior, demonstrating its practical applicability to various black-box reward alignment.
Comments25 pages, 13 figures