arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11886cs.CV

读回:预训练的多模态语言模型是文本到图像生成的零样本奖励模型

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

  • The University of Hong Kong(香港大学)
  • ByteDance Seed(字节跳动Seed)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Runhui Huang, Qihui Zhang, Zhe Liu, Yu Gao, Jie Wu, Hengshuang Zhao

AI总结:

研究提出SpectraReward将预训练多模态语言模型转为图像生成强化学习奖励模型,用图像条件提示对数似然作奖励,还引入Self-SpectraReward形成闭环框架。经广泛实验验证,二者能提升生成性能,表明奖励-策略对齐是关键。

AI中文摘要:

在本文中,我们提出了SpectraReward,一种无需训练的奖励函数,它将预训练的多模态语言模型转变为用于图像生成强化学习的现成奖励模型。SpectraReward不是要求多模态语言模型判断生成的图像或回答分解的验证问题,而是通过单次图像条件下的教师强制前向传递来衡量从生成的图像中恢复原始提示的程度。我们使用平均图像条件提示对数似然作为奖励,直接重用多模态语言模型的预训练图像-文本对齐能力,无需偏好标签和奖励模型微调。我们进一步引入了Self-SpectraReward,这是统一多模态模型的一种特殊情况,其中策略自身的理解分支作为其生成分支的奖励模型,形成了一个无需外部奖励模型或外部知识的闭环自我改进框架。广泛的实验通过涵盖两个扩散模型、三种强化学习算法、来自四个多模态语言模型家族的九个奖励多模态语言模型主干(参数跨度从4B到235B)以及五个分布外文本到图像基准的广泛图像生成强化学习研究验证了SpectraReward。结果表明,SpectraReward和Self-SpectraReward都显著且持续地提高了生成性能,并且优于先前基于多模态语言模型的奖励训练方法。进一步的分析表明,更大的奖励多模态语言模型并不总是更好,而Self-SpectraReward可以匹配或超过大得多的外部奖励模型,这表明奖励-策略对齐是有效图像生成强化学习的关键因素。

英文摘要:

In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single image-conditioned, teacher-forced forward pass. We use the average image-conditioned prompt log-likelihood as the reward, directly reusing the MLLM's pretrained image-text alignment ability without preference labels, reward-model fine-tuning. We further introduce Self-SpectraReward, a special case for unified multimodal models where the policy's own understanding branch serves as the reward model for its generation branch, forming a closed-loop self-improving framework without external reward models or external knowledge. Extensive experiments validate SpectraReward through a broad image-generation RL study covering two diffusion models, three RL algorithms, nine reward MLLM backbones from four MLLM families spanning 4B to 235B parameters, and five out-of-distribution text-to-image benchmarks. Results show that both SpectraReward and Self-SpectraReward significantly and consistently improve generation performance and outperform prior MLLM-derived reward training methods. Further analysis reveals that larger reward MLLMs are not always better, while Self-SpectraReward can match or surpass much larger external reward models, suggesting that reward-policy alignment is a key factor for effective image-generation RL. Project Page: https://huangrh99.github.io/SpectraReward/

↑