arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.05922cs.CV

思考,然后评分:视频奖励建模中的解耦推理与评分

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling

Yuan Wang, Ouxiang Li, Yulong Xu, Borui Liao, Jiajun Liang, Jinghan Li, Meng Wang, Xintao Wang, Pengfei Wan, Kuien Liu, Xiang Wang

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出DeScore模型,通过解耦推理与评分机制提升视频奖励建模的泛化能力,结合冷启动和强化学习优化,实现更稳定的性能。

中文摘要 AI 辅助

近期生成视频模型的进步越来越依赖于训练后和测试时的扩展,这两者都严重依赖于视频奖励模型(RMs)的质量。理想的奖励模型应能准确预测奖励,以符合人类偏好。然而,现有方法面临根本性困境:判别性RMs直接在多模态大语言模型(MLLMs)提取的特征上回归奖励,缺乏显式推理,易导致捷径学习,且依赖大量数据扩展才能泛化。相比之下,生成性RMs结合链式推理(CoT)具有更好的可解释性和泛化潜力,但因推理与评分耦合导致优化瓶颈。为利用CoT推理的泛化优势并缓解耦合推理与评分的训练不稳定性,我们引入DeScore,一种高效且可泛化的视频奖励模型。DeScore采用解耦的『思考-然后评分』范式:MLLM首先生成显式CoT,随后由专用判别性评分模块预测最终奖励。DeScore通过两阶段框架优化:(1)判别性冷启动结合随机掩码机制确保稳健评分能力;(2)双目标强化学习阶段独立优化CoT推理质量和校准最终奖励,确保高质量推理直接转化为更优模型性能。

英文摘要

Recent advances in generative video models are increasingly driven by post-training and test-time scaling, both of which critically depend on the quality of video reward models (RMs). An ideal reward model should predict accurate rewards that align with human preferences across diverse scenarios. However, existing paradigms face a fundamental dilemma: \textit{Discriminative RMs} regress rewards directly on features extracted by multimodal large language models (MLLMs) without explicit reasoning, making them prone to shortcut learning and heavily reliant on massive data scaling for generalization. In contrast, \textit{Generative RMs} with Chain-of-Thought (CoT) reasoning exhibit superior interpretability and generalization potential, as they leverage fine-grained semantic supervision to internalize the rationales behind human preferences. However, they suffer from inherent optimization bottlenecks due to the coupling of reasoning and scoring within a single autoregressive inference chain. To harness the generalization benefits of CoT reasoning while mitigating the training instability of coupled reasoning and scoring, we introduce DeScore, a training-efficient and generalizable video reward model. DeScore employs a decoupled ``think-then-score'' paradigm: an MLLM first generates an explicit CoT, followed by a dedicated discriminative scoring module consisting of a learnable query token and a regression head that predicts the final reward. DeScore is optimized via a two-stage framework: (1) a discriminative cold start incorporating a random mask mechanism to ensure robust scoring capabilities, and (2) a dual-objective reinforcement learning stage that independently refines CoT reasoning quality and calibrates the final reward, ensuring that higher-quality reasoning directly translates to superior model performance.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Kling Team, Kuaishou Technology(快手科技 Kling 团队)
  • Institute of Software Chinese Academy of Sciences(中国科学院软件研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑