VlogReward:学习用于Vlog编辑的多维评估
VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
浏览论文内容
中文总结 AI 辅助
针对Vlog评估主观性强且缺乏标准等问题,提出VlogReward模型。通过专业指导定义评估框架,构建数据集和基准,增强GRPO框架。该模型能提供多维分数与反馈,优于现有MLLMs,助力Vlog创作者及自动化评估完善系统。
中文摘要 AI 辅助
随着Vlog作为一种个性化叙事媒介的迅速兴起,对评估和完善Vlog编辑计划的自动化系统产生了需求。然而,Vlog评估具有高度主观性,由于缺乏标准化标准、数据集和基准以及有效的奖励模型,仍然具有挑战性。为应对这些挑战,我们定义了一个由专业Vlog创作者和产品经理指导的综合Vlog评估框架,建立了六个关键维度的分类法。随后,我们策划了一个包含10万个Vlog编辑的大规模数据集和一个专用基准VRMBench,以评估多模态大语言模型(MLLMs)的Vlog奖励能力。最后,我们提出了VlogReward,一个强大的Vlog奖励模型,它可以提供细粒度的多维分数和可操作的反馈用于迭代改进。从技术上讲,我们通过引入可调整的组间比较奖励来增强组相对策略优化(GRPO)框架,减轻了标准GRPO的“方向盲目性”问题,使模型能够更好地区分不同质量的编辑。VlogReward取得了显著优于现有MLLMs(包括GPT-5和Gemini-3-Pro)的最新结果。我们希望我们的研究能够帮助Vlog创作者并促进自动化的Vlog评估和完善系统。
英文摘要
The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is highly subjective and remains challenging due to a lack of standardized criteria, dataset and benchmark, and effective reward models. To address these challenges, we define a comprehensive vlog evaluation framework guided by professional vlog creators and product managers, establishing a taxonomy of six key dimensions, i.e., Creativity, Consistency, Concept Design, Cinematography, Narration, and Pacing. Subsequently, we curate a large-scale dataset of 100k vlog edits and a dedicated benchmark, VRMBench, to evaluate the vlog rewarding capabilities of Multimodal Large Language Models (MLLMs). Finally, we present VlogReward, a robust vlog reward model that can provide both fine-grained multi-dimensional scores and actionable feedback for iterative refinement. Technically, we enhance the Group Relative Policy Optimization (GRPO) framework by introducing an adjustable inter-group comparison reward, which mitigates the "direction blindness" issue of standard GRPO and enables the model to better distinguish varied-quality edits. VlogReward achieves state-of-the-art results that significantly outperform existing MLLMs, including GPT-5 and Gemini-3-Pro. We hope that our study can help vlog creators and foster automated vlog evaluation and refinement systems.