arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysVista:通过感知-推理-评估循环基准测试视觉语言模型中的物理智能

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

Xinge Peng, Yiting Lu, Tianwu Zhi, Wen Wen, Jianzhao Liu, Xin Li, Zhibo Chen

arXiv 2610.00559首次发表:更新:

发表机构

University of Science and Technology of China; ByteDance; City University of Hong Kong(中国科学技术大学; 字节跳动; 香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhysVista通过闭环认知框架联合评估视觉语言模型的物理状态感知、动态推理与合理性判断,实验发现其在物理推理与评估上存在显著局限,揭示了视觉识别与真实物理理解之间的差距。

AI 中文摘要

视觉语言模型(VLMs)已展现出强大的多模态推理能力,然而它们是否真正捕捉到了现实世界动态背后的物理一致性仍不清楚。现有的基准测试范式往往存在评估碎片化的问题,侧重于孤立的认知阶段,而忽视了感知、推理和物理判断之间的内在协同作用。缺乏整体视角限制了诊断VLMs能否可靠评估新兴生成模型物理真实性的能力。为解决这些问题,我们引入了PhysVista,这是一个通过受人类“看-推理-评估”过程启发的闭环认知框架来评估VLMs物理智能的基准。PhysVista通过联合评估物理状态感知、物理动态推理和物理合理性评估来还原这一循环。它进一步区分了事件级推理和尺度级推理,以实现对物理理解的细粒度分析。此外,PhysVista整合了真实世界和AI生成的视频,允许在多样化的领域和新兴生成场景中进行评估。在多种VLMs上进行的大量实验揭示了物理推理和合理性评估方面的显著局限性,凸显了视觉识别与真正物理理解之间持续存在的差距,并为更具原则性的物理基础多模态智能设计指明了方向。

英文摘要

Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the physical consistency underlying real-world dynamics remains unclear. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inherent synergy between perception, reasoning, and physical judgment. The lack of a holistic perspective limits the ability to diagnose whether VLMs can reliably evaluate the physical authenticity of emerging generative models. To address these issues, we introduce PhysVista, a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process. PhysVista restores this loop by jointly evaluating physical state perception, physical dynamics reasoning, and physical plausibility assessment. It further distinguishes event-level reasoning and scale-level reasoning to enable fine-grained analysis of physical understanding. In addition, PhysVista incorporates both real-world and AI-generated videos, allowing evaluation across diverse domains and emerging generative scenarios. Extensive experiments across a diverse set of VLMs reveal substantial limitations in physical reasoning and plausibility assessment, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.

CommentsAccepted at NeurIPS 2026 (Main Track)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑