发表机构
National University of Singapore; Nanjing University; Kuaishou Technology; University of Oxford(新加坡国立大学; 南京大学; 快手科技; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出VWG-Bench基准测试和Vid-PRE提示增强器,用于评估并提升视频生成模型的推理能力,揭示了渲染与推理的差距并提供了改进方案。
AI 中文摘要
视频生成技术已发展到能够产生视觉上引人入胜且时间上连贯的结果。然而,这些模型是否能够真正地“用视频思考”——执行符号规则、尊重物理规律并追求有意图的目标——仍然是一个悬而未决的问题。现有基准测试仅部分解决了这一问题,常常将视觉质量与认知正确性混为一谈。我们引入了VWG-Bench(视频世界通才基准测试),这是一个涵盖9个推理维度和38个细粒度任务的综合基准测试。为了实现精确诊断,我们设计了一个三级VLM作为评判者的协议,独立评估视频级别的流畅性、任务级别的规则遵循性以及样本级别的目标实现情况。对领先模型的评估揭示了一个显著的差距:虽然模型在渲染得分上表现强劲,但在逻辑密集和规则约束的任务上却持续失败。为了解决这一问题,我们提出了Vid-PRE(视频提示推理器与增强器),这是一种与模型无关的提示重写器,将推理的认知负担转移给专门的VLM。通过使用纯文本奖励的强化学习进行训练,Vid-PRE能够生成简洁、约束感知的提示,而无需视频级别奖励信号的不稳定性。实验表明,Vid-PRE在多种生成器上无需架构修改即可带来显著的推理改进。VWG-Bench和Vid-PRE共同提供了一种严格的诊断工具和一条通往真正视频思考能力的可扩展路径。所有数据和代码均可在此https URL公开获取。
英文摘要
Video generation has advanced to produce visually compelling and temporally coherent results. Yet, whether these models can genuinely think with video--executing symbolic rules, respecting physical laws, and pursuing intentional goals--remains an open question. Existing benchmarks only partially address this, often conflating visual quality with cognitive correctness. We introduce VWG-Bench (Video World Generalist Benchmark), a comprehensive benchmark spanning 9 reasoning dimensions and 38 fine-grained tasks. To enable precise diagnosis, we design a three-level VLM-as-Judge protocol that independently assesses video-level fluency, task-level rule adherence, and sample-level goal realization. Evaluations of leading models reveal a striking gap: while models achieve strong rendering scores, they consistently fail on logic-heavy and rule-constrained tasks. To address this, we propose Vid-PRE (Video Prompt Reasoner and Enhancer), a model-agnostic prompt rewriter that offloads the cognitive burden of reasoning to a dedicated VLM. Trained via reinforcement learning with purely text-based rewards, Vid-PRE produces concise, constraint-aware prompts without the instability of video-level reward signals. Experiments show that Vid-PRE yields substantial reasoning improvements across multiple generators without architectural modifications. Together, VWG-Bench and Vid-PRE offer a rigorous diagnostic lens and a scalable path toward true think-with-video capabilities. All data and code are publicly available at https://huggingface.co/datasets/KlingTeam/VWG-Bench.
CommentsAccepted to ECCV 2026. 46 pages, 41 figures