arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PreviewDiff:扩散潜空间上的多模态评论引导搜索

PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

Vighnesh Subramaniam, Boris Katz, Brian Cheung, Chun-Liang Li, Tomas Pfister, Yale Song

arXiv 2609.36199首次发表:更新:

发表机构

MIT CSAIL; Google Cloud AI Research; UCSF(麻省理工学院计算机科学与人工智能实验室; 谷歌云人工智能研究部; 加州大学旧金山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PreviewDiff提出无需训练的测试时搜索方法,将扩散采样转为多模态评论引导的潜空间搜索,通过部分预览和反馈分支改进生成,在图像视频基准上优于Best-of-N。

AI 中文摘要

扩散模型可以生成引人注目的图像和视频,但它们仍然难以处理使生成结果忠实于提示的组成细节,例如对象数量、属性绑定、空间关系和时间上扎根的动作。提高提示满足度的一种常见方法是在测试时通过Best-of-N采样投入更多计算,但最终样本的选择是固定的。Best-of-N只能在完成的输出中进行选择,无法在 promising 轨迹失败之前修复它。我们引入了PreviewDiff,一种无需训练的测试时搜索方法,将扩散采样从标量搜索转变为对中间潜变量的多模态评论引导搜索。在选定的去噪检查点,PreviewDiff解码部分预览,请求多模态评判者对其进行评分和评论,并利用生成的自然语言反馈在语义提示编辑和局部重新加噪的潜变量延续上进行分支。然后对这些分支进行评分并选择性地向前推进,从而在样本仍可编辑时让验证器计算引导生成。在图像和视频生成基准测试中,PreviewDiff始终优于预算匹配的Best-of-N选择和强标量搜索基线。消融实验表明,早期干预和增加搜索宽度提供了最大的收益,而更深的搜索和额外的语义变体提供了互补的改进。PreviewDiff证明了多模态反馈不仅作为最终验证器,而且作为去噪过程中的主动控制器最为有用。

英文摘要

Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. At selected denoising checkpoints, PreviewDiff decodes a partial preview, asks a multimodal judge to score and critique it, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations. These branches are then scored and selectively rolled forward, allowing verifier compute to guide generation while the sample is still editable. Across image and video generation benchmarks, PreviewDiff consistently improves over budget-matched Best-of-N selection and strong scalar-search baselines. Ablations show that earlier interventions and increased search width provide the largest gains, while deeper search and additional semantic variants offer complementary improvements. PreviewDiff demonstrates that multimodal feedback is most useful not only as a final verifier, but as an active controller inside the denoising process.

Comments24 pages, 11 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑