发表机构
University of Illinois Urbana Champaign; Tsinghua University; University of Waterloo; Massachusetts Institute of Technology; University of British Columbia; Vector Institute; Microsoft Research; Etude AI(伊利诺伊大学厄巴纳-香槟分校; 清华大学; 滑铁卢大学; 麻省理工学院; 不列颠哥伦比亚大学; 矢量研究所; 微软研究院; Etude AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出VGI-bench基准,含27项任务810个实例,评估视频生成模型视觉推理能力,发现最强模型Seedance~2.0仅达51.0%,将推动下一代视频生成模型发展。
AI 中文摘要
近期研究表明,视频生成模型可通过生成帧展现出某些形式的零样本视觉推理能力,但可靠的评估仍具挑战性:基准测试应采用与当前视频模型视觉先验相符的输入,要求有效的演化过程而非仅合理的最终状态,还需校准任务难度以保持挑战性且部分可实现。为此,我们推出VGI-bench,包含27项任务和810个实例,按任务领域和技能标签的两级分类组织,用于对视频生成模型的视觉推理能力进行细粒度评估。我们的评估显示,当前生成系统可解决部分视觉基础推理任务,但仍远未达到可靠水平,即使是最强模型Seedance~2.0,在我们的评估标准下也仅达到51.0%。我们的分析进一步探究了输出失败模式、输入条件敏感性、合成微调的性能迁移边界,以及内部去噪视角揭示的有限自校正能力——后续步骤主要是优化早期假设而非纠正推理错误。我们希望VGI-bench将助力推动下一代视频生成模型的发展,我们将发布代码和数据。
英文摘要
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance 2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. Website: https://hexuan21.github.io/VGI-Bench/