arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

统一模型中原生反思的交叉强化学习

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Yijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu

arXiv 2609.35767首次发表:更新:

发表机构

Nanyang Technological University; Shanghai Jiao Tong University; The University of Tokyo(南洋理工大学; 上海交通大学; 东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出UMM-Reflection,在统一多模态模型中用交叉强化学习联合优化反思文本与图像生成,显著提升多基准图像生成修复能力。

AI 中文摘要

统一多模态模型既能查看图像也能渲染图像,因此原则上它们可以修复自身的生成结果:诊断图像哪里出错,进行修订,观察结果,再诊断。修订是否有益只有在渲染之后才能知晓,因此反思文本和图像生成必须在整个循环中联合学习。对反思轨迹进行监督微调(SFT)提供了一个冷启动,但未能找到高成功率的修复路径,而仅优化渲染器或仅优化单个头的朴素强化学习则留下了大部分增益未被利用。我们引入了UMM-Reflection,该方法在一个统一模型内应用强化学习(RL)来完善反思轨迹:兄弟轨迹共享一个初始图像,因此组相对优势比较反思策略,而一个轨迹级优势同时更新反思标记和基于流的修订,避免了逐轮信用分配的组合爆炸。与单轮编辑或带有外部批评者的流水线不同,信用在轮次之间流动,并流向同一模型的两个角色,且在推理时无需验证器。在BAGEL上,UMM-Reflection相比SFT将GenEval提高了12.05个百分点,并且增益迁移到WISE(+10.97)、OneIG-Bench(+3.48)和T2I-CompBench++(+4.63),这些均未在训练中使用。

英文摘要

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑