发表机构
HDU; CASIA; PKU; SenseTime; FDU; NWPU; NTU; THU(杭州电子科技大学; 中国科学院自动化研究所; 北京大学; 商汤科技; 复旦大学; 西北工业大学; 南洋理工大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出思考奖励模型(TRM),通过先确定案例关键评估点再评分,并引入PD-GRPO优化,在图像生成与编辑基准上达到开源最优,且有效提升多种视觉生成模型。
AI 中文摘要
视觉奖励模型对于评估和改进视觉生成模型至关重要,然而现有方法通常直接将任务条件和候选输出映射为标量奖励,忽略了每个具体案例中应当评估哪些方面。我们提出了“先思考再评分”这一范式,在判断候选表现之前,明确确定每个案例的关键评估点。遵循这一原则,我们提出了思考奖励模型(TRM),该模型制定案例自适应的评分标准,进行基于评分标准的评估,并产生细粒度的逐点奖励。我们进一步观察到,传统的成对偏好优化可能导致评分极化,因此引入了成对双组相对策略优化(PD-GRPO),利用成对监督来提高奖励区分度,同时保留细粒度的逐点评分。在图像生成和编辑奖励建模基准上的大量实验表明,TRM在开源奖励模型中达到了最先进的性能,同时与专有替代方案保持高度竞争力。此外,将TRM作为强化学习的奖励,持续改善了多种视觉生成模型,证明其细粒度、案例自适应的奖励能够转化为有效的视觉生成优化信号。
英文摘要
Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. We further observe that conventional pairwise preference optimization can induce score polarization, and introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing reward-modeling benchmarks demonstrate that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives. Moreover, using TRM as a reward for reinforcement learning consistently improves diverse visual generation models, demonstrating that its fine-grained, case-adaptive rewards translate into effective optimization signals for visual generation.
Comments31 pages