arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2503.23452cs.CV

VideoGen-Eval:基于智能体的视频生成评估系统

VideoGen-Eval: Agent-based System for Video Generation Evaluation

发表机构中国科学技术大学 · 上海交通大学 · 北京大学软件学院
另 2 家 · 查看机构详情
  • USTC(中国科学技术大学)
  • SJTU(上海交通大学)
  • PKUSZ(北京大学软件学院)
  • Tencent(腾讯)
  • BFA(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuhang Yang, Ke Fan, Shangkun Sun, Hongxiang Li, Ailing Zeng, FeiLin Han, Wei Zhai, Wei Liu, Yang Cao, Zheng-Jun Zha

首次发表 更新
浏览论文内容

中文总结 AI 辅助

针对现有视频生成评估系统在提示词、OOD处理及人类偏好对齐上的不足,提出基于智能体的VideoGen-Eval评估系统及含700提示词与12000视频的基准,实验证明其与人类偏好高度一致。

中文摘要 AI 辅助

视频生成的快速发展使得现有评估系统无法评估最先进的模型,主要原因是简单的提示词无法展示模型的能力、固定的评估算子难以处理分布外(OOD)情况,以及计算指标与人类偏好不一致。为弥补这一差距,我们提出VideoGen-Eval,一个集成基于LLM的内容结构化、基于MLLM的内容判断以及针对时间密集维度的补丁工具的智能体评估系统,以实现动态、灵活且可扩展的视频生成评估。此外,我们引入一个视频生成基准来评估现有前沿模型并验证我们评估系统的有效性。它包含700个结构化、内容丰富的提示词(T2V和I2V)以及由20+个模型生成的超过12,000个视频,其中8个前沿模型被选作智能体和人类的定量评估。大量实验验证了我们提出的基于智能体的评估系统与人类偏好高度一致,能可靠完成评估,且基准具有多样性和丰富性。

英文摘要

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators struggling with Out-of-Distribution (OOD) cases, and misalignment between computed metrics and human preferences. To bridge the gap, we propose VideoGen-Eval, an agent evaluation system that integrates LLM-based content structuring, MLLM-based content judgment, and patch tools designed for temporal-dense dimensions, to achieve a dynamic, flexible, and expandable video generation evaluation. Additionally, we introduce a video generation benchmark to evaluate existing cutting-edge models and verify the effectiveness of our evaluation system. It comprises 700 structured, content-rich prompts (both T2V and I2V) and over 12,000 videos generated by 20+ models, among them, 8 cutting-edge models are selected as quantitative evaluation for the agent and human. Extensive experiments validate that our proposed agent-based evaluation system demonstrates strong alignment with human preferences and reliably completes the evaluation, as well as the diversity and richness of the benchmark.

补充信息

↑