发表机构
Ant Group; Beijing Film Academy; Beijing Institute for General Artificial Intelligence (BIGAI)(蚂蚁集团; 北京电影学院; 北京通用人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出扩展的视频评估基准VGA-BenchV2,构建混合评估器架构,并将其作为奖励模型优化视频生成器,实现从基准到模型优化的闭环,提升视频的美学质量与人类偏好对齐性。
AI 中文摘要
我们提出VGA-BenchV2,这是一个扩展的、与人类判断对齐的基准及优化框架,用于联合评估和提升视频生成质量与美学价值。该框架基于VGA-Bench构建,保留了原有的细粒度分类体系,包含美学(Aesthetic)和生成(Generation)两个主要维度,共52个子维度。在该分类体系的指导下,我们整理了1016个多样化提示词,并收集了12种主流视频生成模型生成的超过60000个视频。更重要的是,VGA-BenchV2通过新增36000个任务级标注大幅扩展了人类标注的监督数据,其中包括16200个美学质量标注、13200个美学标签标注和6600个生成质量标注,分别对应VGA-Bench的13.46倍、11.15倍和1.55倍的规模提升。利用这个扩大的标注语料库,我们开发了一种混合评估器架构,包含用于连续美学评分的VAQA-Net,以及两个基于Qwen的大型视觉语言模型评估器VTag-Net和VGQA-Net,分别用于美学标签标注和生成质量评估。大量实验表明,该架构在不同生成模型上与人类判断具有很强的对齐性。除了评估功能外,VGA-BenchV2还进一步引入了从评估到优化的流水线,其中学习到的美学评估器作为奖励模型,用于基于强化学习的生成器微调。这形成了从基准构建、人类监督到自动评估和模型优化的闭环,使视频生成器不仅能在真实性上提升,还能在美学质量和人类偏好对齐上得到改善。相关资源可在该https URL获取。
英文摘要
We introduce VGA-BenchV2, an extended human-aligned benchmark and optimization framework for jointly evaluating and improving video generation quality and aesthetic value. Built upon VGA-Bench, VGA-BenchV2 preserves the original fine-grained taxonomy with two primary dimensions-Aesthetic and Generation-and 52 sub-dimensions. Guided by this taxonomy, we curate 1,016 diverse prompts and collect over 60,000 videos generated by 12 mainstream video generation models. More importantly, VGA-BenchV2 substantially expands human-labeled supervision by adding 36,000 task-level annotations, including 16,200 for aesthetic quality, 13,200 for aesthetic tagging, and 6,600 for generation quality, corresponding to 13.46x, 11.15x, and 1.55x scale-ups over VGA-Bench, respectively. Leveraging this enlarged annotation corpus, we develop a hybrid evaluator architecture consisting of VAQA-Net for continuous aesthetic scoring and two Qwen-based Large Vision-Language Model evaluators, VTag-Net and VGQA-Net, for aesthetic tagging and generation quality assessment. Extensive experiments demonstrate strong alignment with human judgments across diverse generation models. Beyond evaluation, VGA-BenchV2 further introduces an evaluation-to-optimization pipeline, where the learned aesthetic evaluator serves as a reward model for reinforcement learning-based generator fine-tuning. This closes the loop from benchmark construction and human supervision to automated evaluation and model optimization, enabling video generators to improve not only in realism but also in aesthetic quality and human preference alignment. Resources are available at https://huggingface.co/datasets/BestiVictoryLab/VGA-Bench.
CommentsIJCAI 2026