arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VGIF分数:视频生成中时空指令跟随的可解释性和诊断性评估

VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation

Songyu Xu, Xin Wang, Qiang Chen, Xinran Wang, Muxi Diao, Yuxuan Zhang, Kongming Liang, Rui Lin, Zhanyu Ma

arXiv 2607.13527首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; China Telecommunications Group Co., Ltd.(北京邮电大学; 中国电信集团有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视频生成模型遵循长指令能力评估不足问题,提出VGIF-Score框架,含客观完成和主观满意度分支,能生成统一分数,在VGIF-Bench基准实验中对视频生成指令跟随提供可靠、可解释及有诊断作用的评估。

AI 中文摘要

近期视频生成模型在视觉保真度上取得显著进展,但遵循长且具有组合性指令的能力评估不足。现有评估协议存在诸多问题,如依赖简短、语义浅显的提示等。为填补这一空白,我们提出VGIF-Score,一个高度自动化且可解释的视频生成指令跟随评估框架。它由客观完成分支和主观满意度分支组成,能生成统一分数。在VGIF-Bench基准上的实验表明,VGIF-Score能对视频生成指令跟随提供可靠、可解释且有诊断作用的评估。

英文摘要

Recent video generation models (VGMs) have made substantial progress in visual fidelity, yet their ability to follow long, compositional instructions remains insufficiently evaluated. Existing evaluation protocols often rely on prompts that are short and semantically shallow, with limited atomic constraints and weak spatio-temporal dependencies. They also frequently depend on costly human evaluation or handcrafted vision pipelines, while providing little diagnostic insight into which instruction constraints succeed or fail. To address this gap, we propose VGIF-Score, a highly automated and interpretable framework for evaluating instruction following in video generation. VGIF-Score consists of two complementary components: an objective completion branch that parses prompts into a Spatio-Temporal Directed Acyclic Graph (ST-DAG) and performs dependency-aware QA with short-circuit diagnostics, and a subjective satisfaction branch that uses instruction-conditioned AutoRubric to assess cinematography, visual purity, motion smoothness, and physics adherence. Together, these components produce a unified score that captures both objective completion and perceptual satisfaction. We instantiate this framework on VGIF-Bench, a benchmark of 223 long, structurally entangled prompts paired with approximately 4.3K fine-grained evaluation items. Experiments on 14 proprietary and open-source VGMs across more than 3K generated videos show that VGIF-Score provides reliable, interpretable, and diagnostically useful evaluation of video generation instruction following. The code will be available at https://github.com/PRIS-CV/VGIF-SCORE.

CommentsAccepted by PRCV2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑