arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26956cs.CV

RubricRM:基于动态评分规则的生成式奖励模型,用于图像生成与编辑

RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing

发表机构北京大学 · 阿里巴巴集团 · 阿里巴巴达摩院
查看机构详情
  • Peking University(北京大学)
  • Alibaba Group(阿里巴巴集团)
  • Alibaba DAMO Academy(阿里巴巴达摩院)

机构由 AI 辅助整理,请以论文原文为准。

Zijian Kan, Wei Wang, Long Luo, Bing Zhao, Xuan Ren, Weixu Qiao, Wenbo Li, Hu Wei, Lin Qu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出RubricRM框架,为文本到图像生成和图像编辑训练专用模型,其性能优于现有专用奖励模型,且与强专有MLLM评估器竞争力相当,可用于图像生成与编辑的奖励建模。

中文摘要 AI 辅助

奖励模型在视觉生成模型的对齐中发挥着关键作用,但现有多数视觉奖励模型采用单一标量评分或依赖固定标准,无法适配不同指令,这限制了其可解释性和任务敏感性,尤其在文本到图像生成和基于指令的图像编辑场景中,不同输入需要不同的评估维度。我们提出RubricRM,这是一种成对生成式奖励建模框架,它首先生成针对特定输入的评分规则,包含评估维度、权重和评分标准,随后应用该评分规则对候选图像进行评分。我们采用两阶段训练流程,为文本到图像生成和图像编辑训练专用的RubricRM模型:监督微调使模型掌握基于评分规则的评分范式,而GRPO通过细粒度维度级奖励进一步提升评分性能。在多个生成和编辑基准上的实验表明,RubricRM的性能优于现有专用奖励模型,且尽管使用更小的骨干网络,仍与强大的专有多模态大语言模型(MLLM)评估器具有竞争力。我们的模型、数据和代码可在该https URL获取。

英文摘要

Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions. This limits both interpretability and task sensitivity, especially for text-to-image generation and instruction-based image editing, where different inputs require different evaluation dimensions. We propose RubricRM, a pairwise generative reward modeling framework that first produces an input-specific rubric with evaluation dimensions, weights, and scoring criteria, and then applies the rubric to score candidate images. We train dedicated RubricRM models for text-to-image generation and image editing using a two-stage training pipeline: supervised fine-tuning teaches the model the rubric-based scoring paradigm, while GRPO further improves scoring through fine-grained dimension-level rewards. Experiments on multiple generation and editing benchmarks show that RubricRM outperforms existing specialized reward models and remains competitive with strong proprietary MLLM judges despite using smaller backbones. Our models, data, and code are available at https://github.com/zijiankan/RubricRM.

补充信息

↑