arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视频生成提示增强器的统一评估

Towards Unified Evaluation of Prompt Enhancers for Video Generation

Yawen Shao, Yubo Zhu, Ziyun Dai, Zixun Fang, Kai Zhu, Zeyinzi Jiang, Yufeng Ai, Siyang Sun, Haolan Xue, Yu Shang, Yuxiang Bao, Zoubin Bi, Jingming Luo, Jie Xiao, Chaojie Mao, Zhehan Kan, Hongchen Luo, Yu Liu, Sheng Zhong, Wei Tong, Xueyang Fu, Yang Cao, Wei Zhai, Zheng-Jun Zha

arXiv 2610.11736首次发表:更新:

发表机构

University of Science and Technology of China; Alibaba Group; Fudan University; Tsinghua University; Northeastern University(中国科学技术大学; 阿里巴巴集团; 复旦大学; 清华大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有视频生成提示增强器(PE)评估成本高、混淆PE与下游生成器质量的问题,推出首个统一基准PEBench及对应评估框架,揭示PE方法的发展趋势,验证其评分与下游效用高度一致。

AI 中文摘要

现代视频生成器能够实现日益复杂的视觉叙事,这使得提示增强器(Prompt Enhancer, PE)成为连接简洁用户指令与多模态参考内容到结构化电影化规划的关键桥梁。然而,现有的PE评估依赖于渲染后的视频,这会产生大量计算成本和人力成本,减缓PE的训练与迭代过程,且将PE的质量与下游生成器的行为混为一谈。为解决这一缺口,我们推出PEBench,这是首个针对文本到视频、图像到视频、参考到视频提示增强的直接PE评估统一基准。它包含1100个经专家验证的案例和1005个视觉资产,涵盖35项细粒度任务,涉及多样的时间、电影化、视听及多参考要求。此外,我们开发了PEBench评估框架,这是一个基于证据的框架,结合了模态感知事实提取与基于 rubric 的评估,涵盖24项标准。我们对代表性开源和闭源PE方法的系统评估显示,正出现从细粒度描述性扩展向保留意图的电影化规划转变的趋势,而字幕重建和前向细化方法分别在电影化覆盖范围和语义保真度或内部连贯性上展现出互补优势。人工验证表明,PEBench的评分与Wan3.0和MiniMax-H3生成的增强提示及下游视频的专家判断高度一致,这表明提示级评估能够可靠反映下游效用。

英文摘要

Modern video generators can realize increasingly complex visual narratives, positioning the prompt enhancer (PE) as a critical bridge from concise user instructions and multimodal references to structured cinematic plans. However, existing PE evaluation relies on rendered videos, imposing substantial computational and human costs, slowing PE training and iteration, and conflating PE quality with downstream generator behavior. To address this gap, we introduce PEBench, the first unified benchmark for direct PE evaluation across text-to-video, image-to-video, and reference-to-video prompt enhancement. It comprises 1,100 expert-verified cases and 1,005 visual assets, spanning 35 fine-grained tasks with diverse temporal, cinematic, audiovisual, and multi-reference requirements. In addition, we develop PEBench evaluation, an evidence-grounded framework that combines modality-aware fact extraction with rubric-based assessment across 24 criteria. Our systematic evaluation of representative open- and closed-source PE methods reveals an emerging shift from fine-grained descriptive expansion toward intent-preserving cinematic planning, while the caption-reconstruction and forward-refinement methods show complementary strengths in cinematic coverage and semantic fidelity or internal coherence, respectively. Human validation shows that PEBench scores align closely with expert judgments of enhanced prompts and downstream videos from Wan3.0 and MiniMax-H3, indicating that prompt-level evaluation reliably reflects downstream utility.

CommentsProject page: https://github.com/yawen-shao/PEBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑