发表机构
Zhejiang University; ZJU-Hangzhou Global Scientific and Technological Innovation Center; Mohamed Bin Zayed University of Artificial Intelligence(浙江大学; 浙江大学杭州国际科技创新中心; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PhyCheck数据集,通过粗粒度、细粒度及诊断子集提升视频大语言模型的物理规律理解,实验证实其训练效果,同时指出当前模型难以整合因果条件的问题。
AI 中文摘要
具身智能与世界模型要求视频理解系统超越对物体和动作的识别,建立对物理规律的理解。然而,尽管当前视频-语言模型在通用视频理解任务上表现强劲,仍难以可靠判断观测事件是否符合特定物理规律。现有基准主要评估生成视频的物理质量,对系统评估和改进视频大语言模型(VideoLLM)的物理规律理解支持有限。为解决这一缺口,本文提出PhyCheck,一个按两个互补粒度级别组织的视频问答数据集:粗粒度子集要求模型判断视频中呈现的现象是否符合或违反物理规律;细粒度子集进一步检验模型是否能捕捉到导致违反或符合的物理细节。我们将这些子集作为结构化监督以提升物理理解。此外,该数据集包含带有外部因果上下文的诊断子集,可揭示影响物理合理性的隐藏因素,评估模型能否相应调整判断。对微调后的Qwen2.5-VL的实验表明,使用所提数据训练可显著提升物理一致性理解,而诊断子集的评估显示,当前模型仍难以将额外因果条件纳入其决策。这些发现凸显了识别表面不一致与理解潜在物理机制之间的差距,并为评估和改进视频大语言模型的物理理解奠定了基础。
英文摘要
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understanding tasks, current video-language models still struggle to reliably determine whether an observed event conforms to specific physical laws. Existing benchmarks primarily assess the physical quality of generated videos, providing limited support for systematically evaluating and improving the physical-law understanding of Video Large Language Models (VideoLLMs). To address this gap, we introduce PhyCheck, a video question answering dataset organized at two complementary levels of granularity. The coarse-grained subset asks models to determine whether the phenomenon shown in a video conforms to or violates physical laws, while the fine-grained subset further examines whether models can capture physical details responsible for the violation or compliance. We use these subsets as structured supervision to improve physical understanding. In addition, the dataset contains a diagnostic subset with external causal context that reveal hidden factors affecting physical plausibility, assessing whether models can recalibrate their judgments accordingly. Experiments with Fine-tune Qwen2.5-VL show that training with the proposed data substantially improves the understanding of physical-consistency, while evaluations in the diagnostic subset reveal that current models still have difficulty incorporating additional causal conditions into their decisions. These findings highlight the gap between recognizing surface-level inconsistencies and understanding underlying physical mechanisms, and provide a foundation for evaluating and improving physical understanding in Video-LLMs.
Comments15pages, 4 figures, 4 tables