发表机构
Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; University of Virginia; National University of Singapore; University of Michigan, Ann Arbor; The Chinese University of Hong Kong, Shenzhen(中国科学院自动化研究所; 中国科学院大学; 弗吉尼亚大学; 新加坡国立大学; 密歇根大学安娜堡分校; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有视频理解基准未衡量提示词可恢复性的空白,提出VI-Bench基准,基于海量真实提示词和人工验证视频,评估18个视觉语言模型在多种难度下的反演能力,揭示其显著局限。
AI 中文摘要
视频生成领域的最新进展使得基于提示词的控制日益成为AIGC视频生成的核心。提示词规定了视频应描绘的内容及其呈现方式,控制着视觉风格或镜头行为等因素。理解这种可恢复性对于创意复用与编辑,以及评估提示词泄露风险都至关重要。然而,现有的视频理解基准并未衡量这一能力:字幕可能描述了可见内容,但可重放的提示词必须恢复生成视频所需的与生成相关的控制信息。为弥补这一空白,我们提出了VI-Bench,一个基于1610万条真实用户提示词和900个经过人工验证的AIGC视频构建的基准。VI-Bench涵盖三个难度递增的设置,即单次语义锚定、风格与镜头行为控制、以及多次组合反演,并评估五个生成关键维度:主体、动作、场景、风格和镜头。我们使用反演分数(Inversion Score)在VI-Bench上评估了18个具有代表性的视觉语言模型(VLM),包括2个专有模型和16个开源模型,该分数衡量提示词级别上与原始提示词的对齐程度以及重新生成视频的视频级别保真度。结果揭示了显著的局限性:即使最强的模型在反演分数上也仅达到0.632,随着样本需要更丰富的控制和多次推理,性能急剧下降,且模型常常生成看似合理的提示词,但其重新生成的视频与参考视频存在偏差。这些发现表明,视频提示词反演是一种独特且未被充分评估的能力,要求模型将视觉理解转化为可重放的稳定生成控制。
英文摘要
Recent advances in video generation have made prompt-based control increasingly central to AIGC video generation. Prompts specify what a video should depict and how it should be represented, controlling factors such as visual style or camera behavior. Understanding this recoverability is important both for creative reuse and editing, and for assessing prompt leakage risks. However, existing video understanding benchmarks do not measure this capability: a caption may describe what is visible, but a replayable prompt must recover the generation-relevant controls needed to reproduce the video. To address this gap, we introduce VI-Bench, a benchmark built from 16.1 million real-user prompts and 900 human-verified AIGC videos. VI-Bench spans three progressively harder settings, namely single-shot semantic grounding, control over style and camera behavior, and multi-shot compositional inversion, and evaluates five generation-critical dimensions: subject, action, scene, style, and camera. We evaluate 18 representative VLMs, including 2 proprietary and 16 open-source models on VI-Bench, using an Inversion Score that measures prompt-level alignment with the original prompt and video-level fidelity of the regenerated video. The results reveal substantial limitations: even the strongest model achieves only 0.632 on Inversion Score, performance degrades sharply as samples require richer control and multi-shot reasoning, and models often produce plausible prompts whose regenerated videos deviate from the reference. These findings show that video prompt inversion is a distinct and under-evaluated capability requiring models to transform visual understanding into replay-stable generative control.
Comments31 pages