arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RewardVerse:基于评分标准的策略优化用于视频奖励建模

RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

Zhenchen Tang, Yang Li, Songlin Yang, Bo Peng, Xiaotong Zhao, Shuai Li, Haotian Fan, Alan Zhao, Jing Dong

arXiv 2609.22947首次发表:更新:

发表机构

New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; The Hong Kong University of Science and Technology; Tencent(中国科学院自动化研究所模式识别新实验室; 中国科学院大学人工智能学院; 香港科技大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视频奖励模型标量漂移问题,提出基于动态评分标准的RewardVerse框架及两阶段训练算法RGPO,在EvalVerse基准上实现最先进性能,提供稳健可解释的奖励信号。

AI 中文摘要

强化学习(RL)对于优化视频生成模型至关重要,而一个稳健的奖励模型(RM)是其基石。然而,现有的视频奖励模型通常产生不稳定的标量分数,因为它们直接将复杂、主观的视频质量映射为单一分数,而缺乏明确的评估标准。这导致了标量漂移,即评分尺度在不同提示词之间崩溃或偏移,使得奖励在强化学习中不可靠。受专业人工标注工程的启发,我们通过RewardVerse来解决这一问题,这是一个基于评分标准的视频奖励框架,它在评估查询和评分器之间引入了一个动态评分标准作为中间表示。与无约束的直接评分不同,RewardVerse首先生成明确的评估标准,然后执行基于评分标准的评分,提供了一个稳定的语义锚点,从而减轻了标量漂移。为了高效优化这一协作流程,我们提出了评分标准引导的策略优化(RGPO),一种两阶段的训练算法。RGPO首先使用自进化的种子评分标准来预热评分器,然后联合优化评分标准生成器以产生适应查询的评估标准,同时持续使评分器与人类评分对齐。在16维EvalVerse基准和外部数据集上的大量实验表明,RewardVerse减轻了标量漂移,在逐点和成对评估中均达到了最先进的性能,并为视频生成中的强化学习提供了稳健且可解释的奖励信号。

英文摘要

Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable scalar scores because they directly map complex, subjective video quality into a single score without explicit evaluation criteria. This leads to scalar drift, where the scoring scale collapses or shifts across different prompts, making the reward unreliable for RL. Drawing inspiration from professional human annotation engineering, we address this problem with RewardVerse, a rubric-based video reward framework that introduces a dynamic rubric as an intermediate representation between the evaluation query and the scorer. Instead of unconstrained direct scoring, RewardVerse first generates explicit evaluation criteria and then performs rubric-guided scoring, providing a stable semantic anchor that mitigates scalar drift. To efficiently optimize this collaborative pipeline, we propose Rubric-Guided Policy Optimization (RGPO), a two-stage training algorithm. RGPO first warms up the scorer using self-evolving seed rubrics and then jointly optimizes the rubric generator to produce query-adaptive evaluation criteria while continuously aligning the scorer with human ratings. Extensive experiments on the 16-dimensional EvalVerse benchmark and external datasets demonstrate that RewardVerse mitigates scalar drift, achieves state-of-the-art performance on both pointwise and pairwise evaluation, and provides a robust and interpretable reward signal for RL in video generation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑