AI 中文总结
本文提出RAVEN-Eval框架,基于LMM偏好判断与评分准则引导,可自动评估AI视频生成模型,已筛选任务、收集数据、评估模型与LMM并建立排行榜,为AIVGMs评估提供可扩展路径。
AI 中文摘要
AI视频生成技术已取得快速进展并进入广泛商业应用。因此,使用视觉保真度、语义指令遵循等传统评估标准,已越来越难以区分当前最先进AI视频生成模型(AIVGMs)生成视频间的质量差异;与此同时,人工评估需要更多专业知识与持续注意力,大幅提升了标注成本。这一现状要求开发能在最小人工干预下可靠区分先进AIVGMs间细粒度差异的自动评估方法。为应对该挑战,本文提出RAVEN-Eval,一种主要基于“大语言多模态模型作为评判者(LMM-as-a-judge)”范式的、采用评分准则引导的AIVGMs自动评估框架。通过自动任务筛选与质量过滤流水线,RAVEN-Eval筛选出150个文本生成视频(T2V)任务、100个图像生成视频(I2V)任务,并系统收集了超过4500个人工智能生成视频(AIGVs)。RAVEN-Eval的核心是采用评分准则引导的自动LMM偏好判断,其中LMM评判者根据任务特定的评分准则进行成对比较;该框架还引入了基于锚点的模型插入方法,以降低纳入新模型的评估成本。最后,本文评估了20个高性能AIVGMs以及13个LMM评判者的评判能力,并建立了RAVEN-Eval排行榜。总体而言,RAVEN-Eval为快速发展的AIVGMs的自动且可信评估开辟了一条可扩展路径。
英文摘要
AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following. Meanwhile, human evaluation now requires more expertise and sustained attention, substantially increasing annotation costs. This calls for automated evaluation that can reliably distinguish fine-grained differences among advanced AIVGMs with minimal human intervention. To address this challenge, we present RAVEN-Eval, a rubric-guided automated evaluation framework for AIVGMs, built primarily on the LMM-as-a-judge paradigm. Through an automatic task curation and quality-filtering pipeline, RAVEN-Eval curates 150 text-to-video~(T2V) tasks and 100 image-to-video~(I2V) tasks, and systematically collects more than 4,500 AIGVs. At its core, RAVEN-Eval adopts rubric-guided automated LMM preference judgement, in which LMM judges conduct pairwise comparisons according to task-specific rubrics. It further introduces an anchor-based model insertion approach to reduce the evaluation cost of incorporating new models. Finally, we evaluate 20 high-performance AIVGMs, as well as the judging capabilities of 13 LMM judges, and establish the RAVEN-Eval Leaderboards. Overall, RAVEN-Eval paves a scalable path for automatic and trustworthy evaluation of rapidly evolving AIVGMs.