arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24241cs.CVcs.AI

FilmBench:用于电影视频生成的电影级基准测试

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjin… 展开作者

Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing Fan, Mingshuang Tang, Zhong Li, Hengxia Qiang, Weibin Chen, Jinyang Zhen, Bing Zhao, Lin Qu, Jing Li, Hu Wei

首次发表
浏览论文内容

中文总结 AI 辅助

FilmBench是基于电影学院传统专业电影语言开发的文本到视频和参考到视频基准测试。它从获奖影片反向设计提示,依电影分类法评估,开发自动评估代理并开源核心套件,能再现人类模型排名,指出当前模型在动态美学等方面不足。

中文摘要 AI 辅助

视频生成技术的进步不断缩小了人工智能生成的画面与专业制作的画面之间的视觉差距。然而,大多数基准测试仍然从网络来源或语言模型模板中提取提示,并使用未经训练的通用多模态模型进行评分。更根本的是,它们的评估分类仍然很初级,而不是电影实际制作和评判所依据的专业电影语言标准,因此它们评估的是基本视频的合理性,而不是电影级的制作工艺。我们引入了FilmBench,这是一个基于电影学院传统的专业电影语言的文本到视频(T2V)和参考到视频(R2V)基准测试,由北京电影学院和虎鲸数字媒体与娱乐集团电影制片厂的导演和教师共同开发。它基于三个选择。首先,提示是从跨越20种电影类型的获奖影片片段中反向设计的,并由专业导演选择,因此每个提示都锚定在经过验证的真人参考上;提示遵循真实的拍摄列表,并且大多数脚本有多个镜头(1169个提示中有1056个是多镜头的),这与以前的单片段基准测试不同。其次,评估遵循一个三级电影分类法,包括3个轴、12个组件和35个(T2V)+3个(仅R2V)子指标。第三,我们开发了一个内部专家级自动评估代理,并开源了其核心的电影语言操作符套件(FilmOps)。对领先的视频生成模型(9个用于T2V,7个用于R2V)进行基准测试,评估器在模型级斯皮尔曼相关系数为ρ = 0.95(T2V)和0.96(R2V)时再现了人类模型排名。分数远低于以前的网络风格基准测试,在动态美学方面有两个一致的差距,并且从单镜头到多镜头的性能有明显下降,对于较弱的模型来说差距更大。

英文摘要

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.

发表机构

  • Alibaba Group(阿里巴巴集团)
  • Moku Lab, Hujing Digital Media & Entertainment Group(魔氪实验室,虎鲸数字媒体与娱乐集团)
  • Beijing Film Academy(北京电影学院)

机构由 AI 辅助整理,请以论文原文为准。

↑