arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视频中的思考:视频生成器真的能对现实世界进行推理吗?

Thinking in Video: Can Video Generators Really Reason About the Real World?

Yongheng Zhang, Guang Yang, Ruihan Hou, Qiguang Chen, Ziang Liu, Xiaolong Liu, Manman Zhang, Yanchao Hao, Zheng Wei, Hao Wu, Libo Qin, Peishan Dai, Yinghui Li, Di Yin, Xing Sun

arXiv 2607.17523首次发表:更新:

发表机构

Central South University; Tencent; Tsinghua University(中南大学; 腾讯; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

探讨视频生成器能否对现实世界推理,引入因果生成双判断(CGDJ)评估,发现开源模型无明确因果感知却有合理动态,先进闭源系统推理与生成一致性有限,还揭示了视听失调问题。

AI 中文摘要

世界模型和视频生成的最新进展引发了一种新的推理范式,利用视频生成模型来模拟、预测和推理现实世界动态,即“视频中的思考”。但这一设想未经证实,现有指标将感知保真度与语义逻辑分开。为评估视频生成器是否支持此类推理,引入因果生成双判断(CGDJ)从两个角度审核世界模型一致性。将CGDJ应用于代表性生成器发现感知与预测存在差距,开源模型虽无明确因果感知但有合理动态,而先进闭源系统推理与生成的一致性更强但仍有限。进一步分析揭示了视听失调问题。

英文摘要

Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm as Thinking in Video, where video is not merely an output artifact but a medium for constructing, extending, and verifying causal thought. However, this promise remains unverified: convincing rollouts may reflect memorized appearances rather than causal understanding, while existing metrics separate perceptual fidelity from semantic logic. To evaluate whether video generators support such reasoning, we introduce the Causal-Generative Dual-Judge (CGDJ), auditing World Model Consistency from two perspectives. Explicit Causal Perception tests whether a generator reads a video scenario as a reasoning problem through spatio-temporal flattened visual question answering, while Implicit Generative Perception-Prediction Gap evaluates whether it renders the causal consequence as a consistent future video. Applying CGDJ to representative open- and closed-source generators reveals a clear Perception-Prediction Gap: open-source models produce plausible dynamics despite near-zero explicit causal perception, whereas advanced closed-source systems show stronger but still limited alignment between reasoning and generation. Further analysis exposes audio-visual misalignment, where models verbalize correct causal logic more reliably than they render it, challenging the "world simulator" narrative.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑