arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从合成场景到自然提示的可验证视觉奖励迁移

Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts

Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer

arXiv 2609.35641首次发表:更新:

发表机构

University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出可验证视觉奖励(VVR)框架,通过合成场景生成可验证任务,用于强化学习训练,显著提升图像生成模型在指令遵循上的准确率并泛化到自然提示。

AI 中文摘要

在图像生成中精确遵循指令(例如满足物体数量和空间关系)仍然是一个开放的挑战,至少部分原因是它依赖于不可靠的奖励模型(如物体检测器和视觉语言模型)进行学习。我们引入了可验证视觉奖励(VVR),这是第一个用于程序化可验证图像奖励的框架,并表明在其上训练可以泛化到自然提示。每个VVR任务是一个由几何物体及其相互关系组成的场景,我们从中推导出提示和确定性验证器,因此任务可以以任意数量和在任意选定复杂度下生成。我们发布了VVRBench,包含32种约束类型下的10,000个任务,以及VVRBench-Challenge,包含720个更复杂的任务;我们评估的最强模型GPT-Image-2.5在VVRBench-Challenge上解决了21.4%的任务。使用VVR分数作为强化学习的奖励(RLVVR)将Stable Diffusion 3.5 Medium在VVRBench上的准确率从2.8%提高到28.3%,并展示了一致的从易到难的泛化。这些提升扩展到域外基准,将VVR混合到现有目标中进一步提高了整体性能和人类偏好,促使VVR被采用到标准图像生成后训练流程中。

英文摘要

Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR), the first framework for programmatically verifiable image rewards, and show that training on it generalizes to natural prompts. Each VVR task is a scene of geometric objects and relations among them, from which we derive both the prompt and a deterministic verifier, so tasks can be generated in any number and at any chosen complexity. We release VVRBench, with 10,000 tasks over 32 constraint types, and VVRBench-Challenge, with 720 more complex tasks; the strongest model we evaluate---GPT-Image-2.5---solves 21.4% of VVRBench-Challenge. Using VVR scores as rewards for reinforcement learning (RLVVR) raises the accuracy of Stable Diffusion 3.5 Medium on VVRBench from 2.8% to 28.3% and demonstrates consistent easy-to-hard generalization. These gains extend to out-of-domain benchmarks, and mixing VVR into existing objectives further improves overall performance and human preference, motivating the adoption of VVR into standard image generation post-training recipes.

Comments33 pages, 10 figures, 18 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑