VideoScore2:生成式视频评估中评分前先思考
VideoScore2: Think before You Score in Generative Video Evaluation
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- University of Waterloo(滑铁卢大学)
- Independent(独立)
- M-A-P
- University of Toronto(多伦多大学)
- Zhejiang University(浙江大学)
- Abaka AI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对生成式视频评估缺乏可解释多维判断的问题,本文提出VideoScore2,基于VideoFeedback2和SFT+GRPO训练,输出三维评分与思维链,并在域内外基准及Best-of-N奖励建模中表现优异。
AI中文摘要:
文本到视频生成的最新进展产生了越来越逼真和多样化的内容,但由于这些视频具有涵盖视觉质量、语义对齐和物理一致性的多面性,评估此类视频仍是一项根本性挑战。现有评估器和奖励模型局限于单一的不透明分数,缺乏可解释性,或仅提供粗略分析,因此不足以捕捉视频质量评估的综合性质。我们提出 VideoScore2,这是一个多维、可解释且与人类对齐的框架,可显式评估视觉质量、文本到视频对齐以及物理/常识一致性,同时生成详细的思维链理由。我们的模型在包含27,168个带人工标注视频的大规模数据集 VideoFeedback2 上训练,这些视频具有三个维度的分数和推理轨迹;训练采用监督微调,随后使用 Group Relative Policy Optimization(GRPO)进行强化学习的两阶段流程,以增强分析鲁棒性。大量实验表明,VideoScore2 取得了优异性能:在我们的域内基准 VideoScore-Bench-v2 上准确率为44.35(提升5.94),在四个域外基准(VideoGenReward-Bench、VideoPhy2等)上的平均性能为50.37(提升4.32),同时提供可解释评估,通过用于 Best-of-N 采样的有效奖励建模,弥合评估与可控生成之间的差距。项目页面:https://tiger-ai-lab.github.io/VideoScore2/
英文摘要:
Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic alignment, and physical consistency. Existing evaluators and reward models are limited to single opaque scores, lack interpretability, or provide only coarse analysis, making them insufficient for capturing the comprehensive nature of video quality assessment. We present VideoScore2, a multi-dimensional, interpretable, and human-aligned framework that explicitly evaluates visual quality, text-to-video alignment, and physical/common-sense consistency while producing detailed chain-of-thought rationales. Our model is trained on a large-scale dataset VideoFeedback2 containing 27,168 human-annotated videos with both scores and reasoning traces across three dimensions, using a two-stage pipeline of supervised fine-tuning followed by reinforcement learning with Group Relative Policy Optimization (GRPO) to enhance analytical robustness. Extensive experiments demonstrate that VideoScore2 achieves superior performance with 44.35 (+5.94) accuracy on our in-domain benchmark VideoScore-Bench-v2 and 50.37 (+4.32) average performance across four out-of-domain benchmarks (VideoGenReward-Bench, VideoPhy2, etc), while providing interpretable assessments that bridge the gap between evaluation and controllable generation through effective reward modeling for Best-of-N sampling. Project Page: https://tiger-ai-lab.github.io/VideoScore2/